soml-lang 0.0.3 → 0.0.4
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/distribution/edit.d.ts +5 -5
- package/distribution/edit.js +39 -138
- package/distribution/format.d.ts +1 -1
- package/distribution/format.js +1 -1
- package/distribution/index.d.ts +1 -1
- package/distribution/parse.d.ts +3 -3
- package/distribution/parse.js +40 -80
- package/distribution/stringify.d.ts +1 -1
- package/distribution/stringify.js +1 -1
- package/distribution/tree.d.ts +6 -15
- package/distribution/tree.js +6 -22
- package/package.json +1 -1
- package/readme.md +15 -16
package/distribution/edit.d.ts
CHANGED
|
@@ -13,19 +13,19 @@ Change one value in a document, and keep everything else as it is written: comme
|
|
|
13
13
|
The value at `path` is replaced, or added when it does not exist yet, and an `undefined` value removes it. Only the changed part of the text is rewritten, and an edit to a formatted document leaves it formatted.
|
|
14
14
|
|
|
15
15
|
- A new value is written as `stringify()` writes it, at the indentation of its line. So `0xFF` that is replaced by `255n` becomes `255`, and `8080` becomes `8080.0` unless you pass `8080n` or the `integers: 'number'` option. The members of a new object keep their order, unless you pass the `canonical: true` option. In a container that is on one line, it is written on one line too, so `[1, 2]` with a new item `{a: 3n}` becomes `[1, 2, {a: 3}]`.
|
|
16
|
-
- A new member goes after the last member of its object
|
|
16
|
+
- A new member goes after the last member of its object. Missing objects on the way are created.
|
|
17
17
|
- A new item can be added at the end of an array, with the index that is its length.
|
|
18
18
|
- A new member or item goes on its own line without a comma, after the comments that the one before it owns (see below), and the commas of the other members and items stay as they are. When something follows the member or item before it on the same line, such as the closing bracket of the one-line container `{a: 1}`, it goes on that line after a comma, as in `{a: 1, b: 2}`. In an empty `[]` or `{}`, it goes on a line of its own, unless that container is inside a container on one line, so `a: [1, []]` becomes `a: [1, [2]]`.
|
|
19
19
|
- A removed member or item takes the comments it owns, which are the ones that `format()` keeps with it: the comments after it on its line, also after its comma when nothing else follows there, and the block comments before it on its line, after the comma or bracket before it. So removing `2` from `[1, /* note *\/ 2]` gives `[1]`. A comment on a line of its own belongs to no member or item, so it stays. A removed member or item also removes its lines when nothing else is on them. A line that a block comment after it continues onto counts as one of its lines. Removing every item of a container closes it up to `[]` or `{}`, unless a comment is left inside.
|
|
20
20
|
- A removed member or item takes its comma with it. When the removed items are the last ones and have no comma after them, they take the comma directly before them instead, when only spaces, tabs, and the comments they own are between, so `[1, 2]` becomes `[1]`.
|
|
21
|
-
- Removing the
|
|
21
|
+
- Removing the only member of a document without braces leaves `{}`.
|
|
22
22
|
|
|
23
23
|
Comments inside a value that is replaced or removed are removed with it. In a layout that `format()` never writes, an edit can leave odd spacing, such as a new item after a closing block string delimiter on its line that is indented differently from its neighbors. The document is always valid and has the right value. Removing a value that does not exist changes nothing and is not an error: a missing member, a missing object or array on the way, or an index at or past the end of its array. So removing the same path twice is safe, and `edit(text, path, undefined) === text` tells whether something was removed.
|
|
24
24
|
|
|
25
25
|
The result is parsed before it is returned, so a bug in `edit()` throws an `Error` rather than returning a broken document.
|
|
26
26
|
|
|
27
27
|
@param text - The document.
|
|
28
|
-
@param path - The keys and array indexes that lead to the value, such as `['servers', 0, 'port']`.
|
|
28
|
+
@param path - The keys and array indexes that lead to the value, such as `['servers', 0, 'port']`.
|
|
29
29
|
@param value - The new value, of the types that `stringify()` accepts, or `undefined` to remove the value.
|
|
30
30
|
@param options - How an int is represented in `value`, and whether to sort its members, as for `stringify()`.
|
|
31
31
|
@returns The changed document.
|
|
@@ -40,8 +40,8 @@ import {edit} from 'soml-lang';
|
|
|
40
40
|
edit('name: \'api\' # The service\nport: 8080\n', ['port'], 9090n);
|
|
41
41
|
//=> "name: 'api' # The service\nport: 9090\n"
|
|
42
42
|
|
|
43
|
-
edit('postgres
|
|
44
|
-
//=> "postgres
|
|
43
|
+
edit('postgres: {host: \'db\'}\n', ['postgres', 'port'], 5432n);
|
|
44
|
+
//=> "postgres: {host: 'db', port: 5432}\n"
|
|
45
45
|
|
|
46
46
|
edit('a: 1\nb: 2\n', ['a'], undefined);
|
|
47
47
|
//=> 'b: 2\n'
|
package/distribution/edit.js
CHANGED
|
@@ -8,19 +8,19 @@ Change one value in a document, and keep everything else as it is written: comme
|
|
|
8
8
|
The value at `path` is replaced, or added when it does not exist yet, and an `undefined` value removes it. Only the changed part of the text is rewritten, and an edit to a formatted document leaves it formatted.
|
|
9
9
|
|
|
10
10
|
- A new value is written as `stringify()` writes it, at the indentation of its line. So `0xFF` that is replaced by `255n` becomes `255`, and `8080` becomes `8080.0` unless you pass `8080n` or the `integers: 'number'` option. The members of a new object keep their order, unless you pass the `canonical: true` option. In a container that is on one line, it is written on one line too, so `[1, 2]` with a new item `{a: 3n}` becomes `[1, 2, {a: 3}]`.
|
|
11
|
-
- A new member goes after the last member of its object
|
|
11
|
+
- A new member goes after the last member of its object. Missing objects on the way are created.
|
|
12
12
|
- A new item can be added at the end of an array, with the index that is its length.
|
|
13
13
|
- A new member or item goes on its own line without a comma, after the comments that the one before it owns (see below), and the commas of the other members and items stay as they are. When something follows the member or item before it on the same line, such as the closing bracket of the one-line container `{a: 1}`, it goes on that line after a comma, as in `{a: 1, b: 2}`. In an empty `[]` or `{}`, it goes on a line of its own, unless that container is inside a container on one line, so `a: [1, []]` becomes `a: [1, [2]]`.
|
|
14
14
|
- A removed member or item takes the comments it owns, which are the ones that `format()` keeps with it: the comments after it on its line, also after its comma when nothing else follows there, and the block comments before it on its line, after the comma or bracket before it. So removing `2` from `[1, /* note *\/ 2]` gives `[1]`. A comment on a line of its own belongs to no member or item, so it stays. A removed member or item also removes its lines when nothing else is on them. A line that a block comment after it continues onto counts as one of its lines. Removing every item of a container closes it up to `[]` or `{}`, unless a comment is left inside.
|
|
15
15
|
- A removed member or item takes its comma with it. When the removed items are the last ones and have no comma after them, they take the comma directly before them instead, when only spaces, tabs, and the comments they own are between, so `[1, 2]` becomes `[1]`.
|
|
16
|
-
- Removing the
|
|
16
|
+
- Removing the only member of a document without braces leaves `{}`.
|
|
17
17
|
|
|
18
18
|
Comments inside a value that is replaced or removed are removed with it. In a layout that `format()` never writes, an edit can leave odd spacing, such as a new item after a closing block string delimiter on its line that is indented differently from its neighbors. The document is always valid and has the right value. Removing a value that does not exist changes nothing and is not an error: a missing member, a missing object or array on the way, or an index at or past the end of its array. So removing the same path twice is safe, and `edit(text, path, undefined) === text` tells whether something was removed.
|
|
19
19
|
|
|
20
20
|
The result is parsed before it is returned, so a bug in `edit()` throws an `Error` rather than returning a broken document.
|
|
21
21
|
|
|
22
22
|
@param text - The document.
|
|
23
|
-
@param path - The keys and array indexes that lead to the value, such as `['servers', 0, 'port']`.
|
|
23
|
+
@param path - The keys and array indexes that lead to the value, such as `['servers', 0, 'port']`.
|
|
24
24
|
@param value - The new value, of the types that `stringify()` accepts, or `undefined` to remove the value.
|
|
25
25
|
@param options - How an int is represented in `value`, and whether to sort its members, as for `stringify()`.
|
|
26
26
|
@returns The changed document.
|
|
@@ -35,8 +35,8 @@ import {edit} from 'soml-lang';
|
|
|
35
35
|
edit('name: \'api\' # The service\nport: 8080\n', ['port'], 9090n);
|
|
36
36
|
//=> "name: 'api' # The service\nport: 9090\n"
|
|
37
37
|
|
|
38
|
-
edit('postgres
|
|
39
|
-
//=> "postgres
|
|
38
|
+
edit('postgres: {host: \'db\'}\n', ['postgres', 'port'], 5432n);
|
|
39
|
+
//=> "postgres: {host: 'db', port: 5432}\n"
|
|
40
40
|
|
|
41
41
|
edit('a: 1\nb: 2\n', ['a'], undefined);
|
|
42
42
|
//=> 'b: 2\n'
|
|
@@ -143,7 +143,7 @@ class Editor {
|
|
|
143
143
|
this.#descend(element, index + 1, depth + 1);
|
|
144
144
|
}
|
|
145
145
|
else if (value === undefined) {
|
|
146
|
-
this.#
|
|
146
|
+
this.#removeNode(array, element);
|
|
147
147
|
}
|
|
148
148
|
else {
|
|
149
149
|
this.#replace(element, depth + 1, array);
|
|
@@ -151,66 +151,30 @@ class Editor {
|
|
|
151
151
|
}
|
|
152
152
|
#editInObject(object, index, depth) {
|
|
153
153
|
const path = this.#path;
|
|
154
|
-
|
|
154
|
+
const key = path[index];
|
|
155
|
+
if (typeof key !== 'string') {
|
|
155
156
|
throw new TypeError(`Cannot edit ${describePath(path)}, because ${describeParent(path, index)} is an object, so it needs a key, not an index`);
|
|
156
157
|
}
|
|
157
158
|
const value = this.#value;
|
|
158
|
-
const
|
|
159
|
-
|
|
160
|
-
|
|
161
|
-
|
|
162
|
-
|
|
163
|
-
let siblingLength = 0;
|
|
164
|
-
for (const member of object.members) {
|
|
165
|
-
const { segments } = member.key;
|
|
166
|
-
let length = 0;
|
|
167
|
-
while (length < segments.length && length < rest.length && rest[length] === segments[length].value) {
|
|
168
|
-
length++;
|
|
159
|
+
const member = object.members.find(member => member.key.value === key);
|
|
160
|
+
if (member !== undefined) {
|
|
161
|
+
if (index < path.length - 1) {
|
|
162
|
+
this.#isInsideOneLine = this.#isOneLine(object);
|
|
163
|
+
this.#descend(member.value, index + 1, depth + 1);
|
|
169
164
|
}
|
|
170
|
-
if (
|
|
171
|
-
|
|
172
|
-
this.#isInsideOneLine = this.#isOneLine(object);
|
|
173
|
-
this.#descend(member.value, index + length, depth + length);
|
|
174
|
-
}
|
|
175
|
-
else if (value === undefined) {
|
|
176
|
-
this.#removeMembers(object, [member], length);
|
|
177
|
-
}
|
|
178
|
-
else {
|
|
179
|
-
this.#replaceMemberValue(member, depth + length, object);
|
|
180
|
-
}
|
|
181
|
-
return;
|
|
182
|
-
}
|
|
183
|
-
if (length === rest.length) {
|
|
184
|
-
dottedMembers.push(member);
|
|
185
|
-
}
|
|
186
|
-
else if (length > 0 && length >= siblingLength) {
|
|
187
|
-
sibling = member;
|
|
188
|
-
siblingLength = length;
|
|
189
|
-
}
|
|
190
|
-
}
|
|
191
|
-
// An index where dotted keys built an object, as `['a', 0]` for `a.b: 1`, matches no key segment.
|
|
192
|
-
if (sibling !== undefined && typeof rest[siblingLength] !== 'string') {
|
|
193
|
-
throw new TypeError(`Cannot edit ${describePath(path)}, because ${describeParent(path, index + siblingLength)} is an object, so it needs a key, not an index`);
|
|
194
|
-
}
|
|
195
|
-
if (dottedMembers.length > 0) {
|
|
196
|
-
if (value === undefined) {
|
|
197
|
-
this.#removeMembers(object, dottedMembers, rest.length);
|
|
165
|
+
else if (value === undefined) {
|
|
166
|
+
this.#removeMember(object, member);
|
|
198
167
|
}
|
|
199
168
|
else {
|
|
200
|
-
|
|
201
|
-
const [first, ...others] = dottedMembers;
|
|
202
|
-
this.#replaceMembers(object, first, others, `${this.#keyPrefix(first, rest.length)}: ${this.#valueText(value, depth + rest.length, this.#indentation(first.range[0]), object)}`);
|
|
169
|
+
this.#replaceMemberValue(member, depth + 1, object);
|
|
203
170
|
}
|
|
204
171
|
return;
|
|
205
172
|
}
|
|
206
173
|
if (value === undefined) {
|
|
207
174
|
return;
|
|
208
175
|
}
|
|
209
|
-
|
|
210
|
-
|
|
211
|
-
const key = `${sibling === undefined ? '' : `${this.#keyPrefix(sibling, siblingLength)}.`}${formatKey(rest[siblingLength])}`;
|
|
212
|
-
const newValue = nest(path, index + keyLength, value);
|
|
213
|
-
this.#insert(object, sibling ?? object.members.at(-1), indentation => `${key}: ${this.#valueText(newValue, depth + keyLength, indentation, object)}`);
|
|
176
|
+
const newValue = nest(path, index + 1, value);
|
|
177
|
+
this.#insert(object, object.members.at(-1), indentation => `${formatKey(key)}: ${this.#valueText(newValue, depth + 1, indentation, object)}`);
|
|
214
178
|
}
|
|
215
179
|
/*
|
|
216
180
|
Edits the path from `index` on, inside `node`, whose own depth is `depth`.
|
|
@@ -252,93 +216,49 @@ class Editor {
|
|
|
252
216
|
return offset;
|
|
253
217
|
}
|
|
254
218
|
/*
|
|
255
|
-
Removes `
|
|
219
|
+
Removes `member`. Removing the only member of a document without braces leaves `{}` in its place, because a document is never empty, and the comments that the member owned stay.
|
|
256
220
|
*/
|
|
257
|
-
#
|
|
258
|
-
|
|
259
|
-
|
|
260
|
-
const removed = new Set(members);
|
|
261
|
-
const isParentEmptied = object.members.every(member => removed.has(member) || !this.#sharesPrefix(member, first, parentLength));
|
|
262
|
-
if (!isParentEmptied || (parentLength === 0 && object.braced)) {
|
|
263
|
-
this.#removeNodes(object, members);
|
|
221
|
+
#removeMember(object, member) {
|
|
222
|
+
if (object.braced || object.members.length > 1) {
|
|
223
|
+
this.#removeNode(object, member);
|
|
264
224
|
return;
|
|
265
225
|
}
|
|
266
|
-
this.#
|
|
226
|
+
this.#replacements.push({ start: member.range[0], end: member.range[1], text: '{}' });
|
|
267
227
|
}
|
|
268
228
|
/*
|
|
269
|
-
|
|
270
|
-
*/
|
|
271
|
-
#replaceMembers(object, first, others, text) {
|
|
272
|
-
this.#replacements.push({ start: first.range[0], end: first.range[1], text });
|
|
273
|
-
this.#removeNodes(object, others);
|
|
274
|
-
}
|
|
275
|
-
/*
|
|
276
|
-
Removes `nodes` from `container`.
|
|
229
|
+
Removes `node` from `container`.
|
|
277
230
|
|
|
278
|
-
|
|
231
|
+
The removal also takes a blank line that would be left next to another one or directly inside a bracket. When the only item of a braced container goes and no comment is left inside, its brackets close up, as `[]`.
|
|
279
232
|
*/
|
|
280
|
-
#
|
|
233
|
+
#removeNode(container, node) {
|
|
281
234
|
const items = container.type === 'Array' ? container.elements : container.members;
|
|
282
|
-
|
|
283
|
-
|
|
284
|
-
|
|
285
|
-
|
|
286
|
-
|
|
287
|
-
// When the removed items reach the end and the last one has no comma after it, the first of them takes the comma before it, so that no trailing comma is left, as `[1, 2]` becomes `[1]`.
|
|
288
|
-
const lastItem = items.at(-1);
|
|
289
|
-
const firstToTakeCommaBefore = lastItem !== undefined && this.#text.charCodeAt(this.#skipTrivia(lastItem.range[1])) !== COMMA ? items[trailingStart] : undefined;
|
|
290
|
-
const removals = [];
|
|
291
|
-
for (const removal of nodes.flatMap(node => this.#removals(node, node === firstToTakeCommaBefore))) {
|
|
292
|
-
const previous = removals.at(-1);
|
|
293
|
-
if (previous !== undefined && removal.start < previous.end) {
|
|
294
|
-
// A removal that took the comma before it overlaps the one before.
|
|
295
|
-
previous.end = Math.max(previous.end, removal.end);
|
|
296
|
-
previous.isWholeLines = false;
|
|
297
|
-
}
|
|
298
|
-
else if (removal.isWholeLines && previous?.isWholeLines && this.#skipBlankLines(previous.end) === removal.start) {
|
|
299
|
-
previous.end = removal.end;
|
|
300
|
-
}
|
|
301
|
-
else {
|
|
302
|
-
removals.push(removal);
|
|
303
|
-
}
|
|
235
|
+
// When the removed item is the last one and has no comma after it, it takes the comma before it, so that no trailing comma is left, as `[1, 2]` becomes `[1]`.
|
|
236
|
+
const shouldTakeCommaBefore = node === items.at(-1) && this.#text.charCodeAt(this.#skipTrivia(node.range[1])) !== COMMA;
|
|
237
|
+
const removal = this.#nodeRemoval(node, shouldTakeCommaBefore);
|
|
238
|
+
if (removal.isWholeLines) {
|
|
239
|
+
this.#takeBlankLine(removal, container, this.#contentStart(container));
|
|
304
240
|
}
|
|
305
|
-
|
|
306
|
-
const contentStart = this.#contentStart(container);
|
|
307
|
-
for (const removal of removals) {
|
|
308
|
-
if (removal.isWholeLines) {
|
|
309
|
-
this.#takeBlankLine(removal, container, contentStart);
|
|
310
|
-
}
|
|
311
|
-
}
|
|
312
|
-
if (nodes.length === items.length && isBraced(container) && !this.#isCommentLeft(container, removals)) {
|
|
241
|
+
if (items.length === 1 && isBraced(container) && !this.#isCommentLeft(container, removal)) {
|
|
313
242
|
const [start, end] = container.range;
|
|
314
243
|
this.#replacements.push({ start: start + 1, end: end - 1, text: '' });
|
|
315
244
|
return;
|
|
316
245
|
}
|
|
317
|
-
|
|
318
|
-
this.#replacements.push({ start, end, text: '' });
|
|
319
|
-
}
|
|
246
|
+
this.#replacements.push({ start: removal.start, end: removal.end, text: '' });
|
|
320
247
|
}
|
|
321
248
|
/*
|
|
322
|
-
Whether a comment inside `container` is outside
|
|
249
|
+
Whether a comment inside `container` is outside `removal`.
|
|
323
250
|
*/
|
|
324
|
-
#isCommentLeft(container,
|
|
251
|
+
#isCommentLeft(container, removal) {
|
|
325
252
|
const [start, end] = container.range;
|
|
326
|
-
let removalIndex = 0;
|
|
327
253
|
return this.#comments.some(comment => {
|
|
328
254
|
const [commentStart] = comment.range;
|
|
329
|
-
|
|
330
|
-
return false;
|
|
331
|
-
}
|
|
332
|
-
while (removalIndex < removals.length && removals[removalIndex].end <= commentStart) {
|
|
333
|
-
removalIndex++;
|
|
334
|
-
}
|
|
335
|
-
return removalIndex === removals.length || commentStart < removals[removalIndex].start;
|
|
255
|
+
return commentStart > start && commentStart < end && (commentStart < removal.start || commentStart >= removal.end);
|
|
336
256
|
});
|
|
337
257
|
}
|
|
338
258
|
/*
|
|
339
259
|
What removing a member or an item, with the comments it owns and its comma, takes out of the text. With `shouldTakeCommaBefore`, a comma directly before it goes too.
|
|
340
260
|
*/
|
|
341
|
-
#
|
|
261
|
+
#nodeRemoval(node, shouldTakeCommaBefore) {
|
|
342
262
|
const text = this.#text;
|
|
343
263
|
const start = this.#ownedStart(node.range[0]);
|
|
344
264
|
const end = this.#ownedEnd(node.range[1]);
|
|
@@ -355,7 +275,7 @@ class Editor {
|
|
|
355
275
|
if (removalStart !== start && !removal.isWholeLines && this.#commentEnds.has(removal.end)) {
|
|
356
276
|
removal.end = skipSpacesBack(text, removal.end);
|
|
357
277
|
}
|
|
358
|
-
return
|
|
278
|
+
return removal;
|
|
359
279
|
}
|
|
360
280
|
/*
|
|
361
281
|
The start of the block comments that a member or an item at `offset` owns before it: those on its line after the comma or the opening bracket before it, as `/* note *\/` in `[1, /* note *\/ 2]`. The formatter keeps them in front of it. It is `offset` when there are none.
|
|
@@ -431,7 +351,7 @@ class Editor {
|
|
|
431
351
|
return offset === this.#text.length || (isBraced(container) && skipSpaces(this.#text, offset) === container.range[1] - 1);
|
|
432
352
|
}
|
|
433
353
|
/*
|
|
434
|
-
The start of the first line that is not blank after the line that opens `container`, or after the start of the text. The line that opens a container ends after the comments that follow its bracket, and when an item follows on it, this is `undefined`.
|
|
354
|
+
The start of the first line that is not blank after the line that opens `container`, or after the start of the text. The line that opens a container ends after the comments that follow its bracket, and when an item follows on it, this is `undefined`.
|
|
435
355
|
*/
|
|
436
356
|
#contentStart(container) {
|
|
437
357
|
if (!isBraced(container)) {
|
|
@@ -511,25 +431,6 @@ class Editor {
|
|
|
511
431
|
});
|
|
512
432
|
}
|
|
513
433
|
/*
|
|
514
|
-
The first `length` segments of a member's key, as they are written.
|
|
515
|
-
*/
|
|
516
|
-
#keyPrefix(member, length) {
|
|
517
|
-
const { segments } = member.key;
|
|
518
|
-
return this.#text.slice(segments[0].range[0], segments[length - 1].range[1]);
|
|
519
|
-
}
|
|
520
|
-
/*
|
|
521
|
-
Whether the first `length` segments of the two keys are the same. A key with fewer segments always differs in one of its own, because a key that is a prefix of another one collides with it.
|
|
522
|
-
*/
|
|
523
|
-
#sharesPrefix(member, other, length) {
|
|
524
|
-
const { segments } = member.key;
|
|
525
|
-
for (let index = 0; index < length; index++) {
|
|
526
|
-
if (segments[index].value !== other.key.segments[index].value) {
|
|
527
|
-
return false;
|
|
528
|
-
}
|
|
529
|
-
}
|
|
530
|
-
return true;
|
|
531
|
-
}
|
|
532
|
-
/*
|
|
533
434
|
The spaces and tabs at the start of the line that holds `offset`.
|
|
534
435
|
*/
|
|
535
436
|
#indentation(offset) {
|
package/distribution/format.d.ts
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
/**
|
|
2
2
|
Format a document. Returns the document with its layout normalized, ending with one line feed.
|
|
3
3
|
|
|
4
|
-
The layout follows the formatter in the specification: one tab per level, every member and item on its own line with no commas, and no trailing whitespace or runs of blank lines. An object or an array whose brackets are on one line stays on one line, as in `ports: [80, 443]`, with a comma and a space between its members or items. To give a one-line container one member or item per line, put a line break anywhere inside it. Comments, member order,
|
|
4
|
+
The layout follows the formatter in the specification: one tab per level, every member and item on its own line with no commas, and no trailing whitespace or runs of blank lines. An object or an array whose brackets are on one line stays on one line, as in `ports: [80, 443]`, with a comma and a space between its members or items. To give a one-line container one member or item per line, put a line break anywhere inside it. Comments, member order, block strings, and the spelling of every value stay as they are, so the value never changes. The one change inside a block comment is the same layout rule: trailing whitespace is removed, and runs of blank lines collapse to one.
|
|
5
5
|
|
|
6
6
|
In detail, as the specification states:
|
|
7
7
|
|
package/distribution/format.js
CHANGED
|
@@ -4,7 +4,7 @@ const INDENT = '\t';
|
|
|
4
4
|
/**
|
|
5
5
|
Format a document. Returns the document with its layout normalized, ending with one line feed.
|
|
6
6
|
|
|
7
|
-
The layout follows the formatter in the specification: one tab per level, every member and item on its own line with no commas, and no trailing whitespace or runs of blank lines. An object or an array whose brackets are on one line stays on one line, as in `ports: [80, 443]`, with a comma and a space between its members or items. To give a one-line container one member or item per line, put a line break anywhere inside it. Comments, member order,
|
|
7
|
+
The layout follows the formatter in the specification: one tab per level, every member and item on its own line with no commas, and no trailing whitespace or runs of blank lines. An object or an array whose brackets are on one line stays on one line, as in `ports: [80, 443]`, with a comma and a space between its members or items. To give a one-line container one member or item per line, put a line break anywhere inside it. Comments, member order, block strings, and the spelling of every value stay as they are, so the value never changes. The one change inside a block comment is the same layout rule: trailing whitespace is removed, and runs of blank lines collapse to one.
|
|
8
8
|
|
|
9
9
|
In detail, as the specification states:
|
|
10
10
|
|
package/distribution/index.d.ts
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
/// <reference lib="esnext.temporal" preserve="true" />
|
|
2
2
|
export { parse, type Value, type ObjectValue, type Document, type ParseOptions, } from './parse.ts';
|
|
3
|
-
export { parseTree, visitorKeys, type Position, type SourceLocation, type Token, type Comment, type DocumentNode, type ObjectNode, type MemberNode, type ArrayNode, type KeyNode, type
|
|
3
|
+
export { parseTree, visitorKeys, type Position, type SourceLocation, type Token, type Comment, type DocumentNode, type ObjectNode, type MemberNode, type ArrayNode, type KeyNode, type StringNode, type IntegerNode, type FloatNode, type BooleanNode, type NullNode, type InstantNode, type DurationNode, type DurationPart, type ValueNode, type Node, } from './tree.ts';
|
|
4
4
|
export { format, formatEdits, type FormatEdit } from './format.ts';
|
|
5
5
|
export { edit, type EditOptions, type PathSegment } from './edit.ts';
|
|
6
6
|
export { stringify, stringifyValue, compareKeys, type StringifyOptions, } from './stringify.ts';
|
package/distribution/parse.d.ts
CHANGED
|
@@ -58,7 +58,7 @@ parse(`
|
|
|
58
58
|
name: 'api-gateway'
|
|
59
59
|
replicas: 3
|
|
60
60
|
timeout: 30.0
|
|
61
|
-
postgres
|
|
61
|
+
postgres: {host: 'db.internal'}
|
|
62
62
|
`);
|
|
63
63
|
//=> {name: 'api-gateway', replicas: 3n, timeout: 30, postgres: {host: 'db.internal'}}
|
|
64
64
|
|
|
@@ -85,7 +85,7 @@ parse(`
|
|
|
85
85
|
name: 'api-gateway'
|
|
86
86
|
replicas: 3
|
|
87
87
|
timeout: 30.0
|
|
88
|
-
postgres
|
|
88
|
+
postgres: {host: 'db.internal'}
|
|
89
89
|
`);
|
|
90
90
|
//=> {name: 'api-gateway', replicas: 3n, timeout: 30, postgres: {host: 'db.internal'}}
|
|
91
91
|
|
|
@@ -112,7 +112,7 @@ parse(`
|
|
|
112
112
|
name: 'api-gateway'
|
|
113
113
|
replicas: 3
|
|
114
114
|
timeout: 30.0
|
|
115
|
-
postgres
|
|
115
|
+
postgres: {host: 'db.internal'}
|
|
116
116
|
`);
|
|
117
117
|
//=> {name: 'api-gateway', replicas: 3n, timeout: 30, postgres: {host: 'db.internal'}}
|
|
118
118
|
|
package/distribution/parse.js
CHANGED
|
@@ -76,7 +76,7 @@ export function parse(text, options) {
|
|
|
76
76
|
return new Parser(source, integers).parseDocument();
|
|
77
77
|
}
|
|
78
78
|
/*
|
|
79
|
-
An instant, as nanoseconds since the Unix epoch, or a duration, as its length in nanoseconds.
|
|
79
|
+
An instant, as nanoseconds since the Unix epoch, or a duration, as its length in nanoseconds.
|
|
80
80
|
*/
|
|
81
81
|
export class Time {
|
|
82
82
|
type;
|
|
@@ -198,26 +198,12 @@ function defineMember(object, key, value) {
|
|
|
198
198
|
object[key] = value;
|
|
199
199
|
}
|
|
200
200
|
}
|
|
201
|
-
function describeValue(value) {
|
|
202
|
-
if (Array.isArray(value)) {
|
|
203
|
-
return 'an array';
|
|
204
|
-
}
|
|
205
|
-
// Every parsed object is created with `{}`, unlike an instant or a duration.
|
|
206
|
-
return value !== null && typeof value === 'object' && Object.getPrototypeOf(value) === Object.prototype ? 'an object' : 'a value';
|
|
207
|
-
}
|
|
208
|
-
function formatPath(path) {
|
|
209
|
-
return path.map(segment => describeKey(segment)).join('.');
|
|
210
|
-
}
|
|
211
201
|
class Parser {
|
|
212
202
|
#source;
|
|
213
203
|
#index = 0;
|
|
214
204
|
#integers;
|
|
215
205
|
#createTime;
|
|
216
206
|
/*
|
|
217
|
-
Objects created by dotted keys. They may be extended by further dotted keys, while an object written with braces is closed.
|
|
218
|
-
*/
|
|
219
|
-
#dottedObjects = new WeakSet();
|
|
220
|
-
/*
|
|
221
207
|
Where the document's collection starts, after the comments and whitespace before it.
|
|
222
208
|
*/
|
|
223
209
|
#documentStart = 0;
|
|
@@ -336,11 +322,11 @@ class Parser {
|
|
|
336
322
|
*/
|
|
337
323
|
#parseEntry(object, depth) {
|
|
338
324
|
const keyStart = this.#index;
|
|
339
|
-
const
|
|
340
|
-
if (depth + path.length - 1 > MAX_DEPTH) {
|
|
341
|
-
this.#failTooDeep(keyStart);
|
|
342
|
-
}
|
|
325
|
+
const key = this.#parseKey();
|
|
343
326
|
const code = this.#code();
|
|
327
|
+
if (code === DOT) {
|
|
328
|
+
this.#failDotInKey(keyStart);
|
|
329
|
+
}
|
|
344
330
|
if (code !== COLON) {
|
|
345
331
|
this.#failMissingColon(keyStart);
|
|
346
332
|
}
|
|
@@ -349,20 +335,23 @@ class Parser {
|
|
|
349
335
|
this.#skipTrivia();
|
|
350
336
|
let value;
|
|
351
337
|
try {
|
|
352
|
-
value = this.#parseValue(depth +
|
|
338
|
+
value = this.#parseValue(depth + 1);
|
|
353
339
|
}
|
|
354
340
|
catch (error) {
|
|
355
341
|
if (error instanceof ParseError) {
|
|
356
|
-
this.#diagnoseBadValue(error,
|
|
342
|
+
this.#diagnoseBadValue(error, key, keyStart, colon);
|
|
357
343
|
}
|
|
358
344
|
throw error;
|
|
359
345
|
}
|
|
360
|
-
|
|
346
|
+
if (Object.hasOwn(object, key)) {
|
|
347
|
+
this.#fail(`Duplicate key ${describeKey(key)}`, keyStart);
|
|
348
|
+
}
|
|
349
|
+
defineMember(object, key, value);
|
|
361
350
|
}
|
|
362
351
|
/*
|
|
363
352
|
Reports a value that failed to parse as what the entry was meant to be, when that is clear.
|
|
364
353
|
*/
|
|
365
|
-
#diagnoseBadValue(error,
|
|
354
|
+
#diagnoseBadValue(error, key, keyStart, colon) {
|
|
366
355
|
// A key that contains a `:`, as in `12:30: 'lunch'`, ends at the first `:`, and the rest is read as the value.
|
|
367
356
|
if (isBareKeyCharacter(this.#code(colon + 1))) {
|
|
368
357
|
this.#diagnoseKeyWithColon(keyStart);
|
|
@@ -372,7 +361,7 @@ class Parser {
|
|
|
372
361
|
const valueStart = this.#index;
|
|
373
362
|
const commentHint = this.#describeCommentAsValue(colon);
|
|
374
363
|
if (isValueOnNextLine) {
|
|
375
|
-
this.#diagnoseMissingValue(valueStart,
|
|
364
|
+
this.#diagnoseMissingValue(valueStart, key, keyStart, commentHint);
|
|
376
365
|
}
|
|
377
366
|
// A value that was left out at the end of the document or of an object.
|
|
378
367
|
const code = this.#code(valueStart);
|
|
@@ -414,7 +403,7 @@ class Parser {
|
|
|
414
403
|
/*
|
|
415
404
|
An entry whose value was left out, as in `a:` followed by `'b': 1` on the next line, reads the next key as the value. That failure is reported as what it is.
|
|
416
405
|
*/
|
|
417
|
-
#diagnoseMissingValue(start,
|
|
406
|
+
#diagnoseMissingValue(start, parent, keyStart, commentHint) {
|
|
418
407
|
// A YAML block sequence.
|
|
419
408
|
if (this.#code(start) === DASH && isSpace(this.#code(start + 1))) {
|
|
420
409
|
this.#fail('Expected a value, but found a “-” list. An array is written in brackets, as in [80, 443]', start);
|
|
@@ -436,22 +425,18 @@ class Parser {
|
|
|
436
425
|
if (firstCode !== SINGLE_QUOTE && firstCode !== DOUBLE_QUOTE && next !== LF && !isSpace(next) && !Number.isNaN(next)) {
|
|
437
426
|
return;
|
|
438
427
|
}
|
|
439
|
-
const hint = commentHint === '' ? this.#describeIndentedKey(
|
|
440
|
-
this.#fail(`Expected a value, but found the key ${
|
|
428
|
+
const hint = commentHint === '' ? this.#describeIndentedKey(parent, key, keyStart, start) : commentHint;
|
|
429
|
+
this.#fail(`Expected a value, but found the key ${describeKey(key)}${hint}`, start);
|
|
441
430
|
}
|
|
442
431
|
/*
|
|
443
432
|
The hint for a key at `start` that is indented under the entry whose value is missing, as YAML nests an object.
|
|
444
433
|
*/
|
|
445
|
-
#describeIndentedKey(
|
|
434
|
+
#describeIndentedKey(parent, key, keyStart, start) {
|
|
446
435
|
const keyIndentation = getIndentation(this.#source, keyStart);
|
|
447
436
|
const indentation = getIndentation(this.#source, start);
|
|
448
|
-
if (keyIndentation === undefined || indentation === undefined || indentation <= keyIndentation) {
|
|
449
|
-
return '';
|
|
450
|
-
}
|
|
451
437
|
// The keys are written as in a document, so that the suggestion is valid, rather than as JSON strings like the rest of the message, whose escapes, such as `\b`, are not all valid.
|
|
452
|
-
const
|
|
453
|
-
|
|
454
|
-
return `. Indentation does not nest objects, so write ${parent}: {${child}: …} or ${parent}.${child}: …`;
|
|
438
|
+
const hint = `. Indentation does not nest objects, so write ${abbreviate(formatKey(parent), 200)}: {${abbreviate(formatKey(key), 200)}: …}`;
|
|
439
|
+
return keyIndentation === undefined || indentation === undefined || indentation <= keyIndentation ? '' : hint;
|
|
455
440
|
}
|
|
456
441
|
#failMissingColon(keyStart) {
|
|
457
442
|
const source = this.#source;
|
|
@@ -468,7 +453,7 @@ class Parser {
|
|
|
468
453
|
}
|
|
469
454
|
}
|
|
470
455
|
const colon = source.indexOf(':', next);
|
|
471
|
-
// A key with a space in it, such as `the name: 1`. The hint is only given for plain words, because quoting a
|
|
456
|
+
// A key with a space in it, such as `the name: 1`. The hint is only given for plain words, because quoting a quoted key would change what it means, and for a `:` that whitespace or the end follows, because quoting the words before the `:` in `server localhost:8080` would give a valid document with another meaning. So `the name:1` gets no hint.
|
|
472
457
|
if (colon !== -1 && next > this.#index && isBareKeyCharacter(nextCode) && colon < findLineEnd(source, next) && isSpaceOrLineEnd(source.charCodeAt(colon + 1))) {
|
|
473
458
|
// Only spaces and tabs are trimmed, because `trimEnd()` would also remove characters that are not whitespace in SOML, such as U+00A0.
|
|
474
459
|
const key = source.slice(keyStart, skipSpacesBack(source, colon));
|
|
@@ -486,14 +471,6 @@ class Parser {
|
|
|
486
471
|
this.#fail(`Expected “:” after the key, but found ${this.#describeHere()}${hint}`);
|
|
487
472
|
}
|
|
488
473
|
#parseKey() {
|
|
489
|
-
const path = [this.#parseKeySegment()];
|
|
490
|
-
while (this.#code() === DOT) {
|
|
491
|
-
this.#index++;
|
|
492
|
-
path.push(this.#parseKeySegment(true));
|
|
493
|
-
}
|
|
494
|
-
return path;
|
|
495
|
-
}
|
|
496
|
-
#parseKeySegment(isAfterDot = false) {
|
|
497
474
|
const source = this.#source;
|
|
498
475
|
const start = this.#index;
|
|
499
476
|
const code = source.charCodeAt(start);
|
|
@@ -512,8 +489,7 @@ class Parser {
|
|
|
512
489
|
if (code === OPEN_BRACKET) {
|
|
513
490
|
this.#diagnoseTableHeader();
|
|
514
491
|
}
|
|
515
|
-
|
|
516
|
-
this.#fail(this.#isAtEnd() ? expected : `${expected}, but found ${this.#describeHere()}${this.#keyQuotingHint()}${this.#slashCommentHint()}`);
|
|
492
|
+
this.#fail(this.#isAtEnd() ? 'Expected a key' : `Expected a key, but found ${this.#describeHere()}${this.#keyQuotingHint()}${this.#slashCommentHint()}`);
|
|
517
493
|
}
|
|
518
494
|
this.#index = end;
|
|
519
495
|
return source.slice(start, end);
|
|
@@ -529,39 +505,23 @@ class Parser {
|
|
|
529
505
|
const key = QUOTABLE_KEY.exec(this.#source.slice(this.#index, this.#index + MAX_DIAGNOSED_LENGTH))?.[0];
|
|
530
506
|
return key === undefined ? KEY_QUOTING_HINT : `${KEY_QUOTING_HINT}, as in '${abbreviate(key)}'`;
|
|
531
507
|
}
|
|
532
|
-
|
|
533
|
-
|
|
534
|
-
|
|
535
|
-
|
|
536
|
-
|
|
537
|
-
|
|
538
|
-
|
|
539
|
-
|
|
540
|
-
defineMember(target, key, child);
|
|
541
|
-
target = child;
|
|
542
|
-
continue;
|
|
543
|
-
}
|
|
544
|
-
// Only an object built by dotted keys is in the set, which is checked next.
|
|
545
|
-
const existing = target[key];
|
|
546
|
-
if (!this.#dottedObjects.has(existing)) {
|
|
547
|
-
const prefix = formatPath(path.slice(0, index + 1));
|
|
548
|
-
const what = describeValue(existing);
|
|
549
|
-
const reason = what === 'an object' ? 'an object written with braces, which is closed' : what;
|
|
550
|
-
this.#fail(`Cannot set ${formatPath(path)}, because ${prefix} is already ${reason} and a dotted key cannot extend it`, keyStart);
|
|
551
|
-
}
|
|
552
|
-
target = existing;
|
|
508
|
+
/*
|
|
509
|
+
A `.` after a key, as in `example.com: 1`. A key is never a path, so the `.` is meant either as part of the key or as nesting.
|
|
510
|
+
*/
|
|
511
|
+
#failDotInKey(keyStart) {
|
|
512
|
+
const source = this.#source;
|
|
513
|
+
let end = this.#index;
|
|
514
|
+
while (end < keyStart + MAX_DIAGNOSED_LENGTH && (isBareKeyCharacter(source.charCodeAt(end)) || source.charCodeAt(end) === DOT)) {
|
|
515
|
+
end++;
|
|
553
516
|
}
|
|
554
|
-
const key =
|
|
555
|
-
|
|
556
|
-
|
|
557
|
-
|
|
558
|
-
}
|
|
559
|
-
this.#fail(`Duplicate key ${formatPath(path)}`, keyStart);
|
|
517
|
+
const key = source.slice(keyStart, end);
|
|
518
|
+
const words = key.split('.');
|
|
519
|
+
// The suggestions are only given for a bare key whose dots each have a word on both sides, so that both are valid.
|
|
520
|
+
if (source.charCodeAt(end) !== COLON || !isBareKeyCharacter(source.charCodeAt(keyStart)) || words.includes('')) {
|
|
521
|
+
this.#fail('A key cannot contain “.” unless it is quoted. Quote the whole key, or use braces to nest, as in a: {b: …}');
|
|
560
522
|
}
|
|
561
|
-
|
|
562
|
-
|
|
563
|
-
#failTooDeep(offset = this.#index) {
|
|
564
|
-
this.#fail(`The document is nested more than ${MAX_DEPTH} levels deep`, offset);
|
|
523
|
+
const nested = `${words.join(': {')}: …${'}'.repeat(words.length - 1)}`;
|
|
524
|
+
this.#fail(`A bare key cannot contain “.”. Quote it, as in '${abbreviate(key)}', or use braces to nest, as in ${abbreviate(nested)}`);
|
|
565
525
|
}
|
|
566
526
|
#parseObject(depth) {
|
|
567
527
|
const object = {};
|
|
@@ -582,7 +542,7 @@ class Parser {
|
|
|
582
542
|
*/
|
|
583
543
|
#parseItems(depth, closing, parseItem) {
|
|
584
544
|
if (depth > MAX_DEPTH) {
|
|
585
|
-
this.#
|
|
545
|
+
this.#fail(`The document is nested more than ${MAX_DEPTH} levels deep`);
|
|
586
546
|
}
|
|
587
547
|
const start = this.#index;
|
|
588
548
|
this.#index++;
|
|
@@ -758,13 +718,13 @@ class Parser {
|
|
|
758
718
|
return;
|
|
759
719
|
}
|
|
760
720
|
const line = source.slice(lineStart, Math.min(findLineEnd(source, this.#index), lineStart + MAX_DIAGNOSED_LENGTH));
|
|
761
|
-
// The name is a valid key, so that the
|
|
762
|
-
const match = /^[\t ]*\[\[?(?<name>[A-Z_a-z][\w\-]*
|
|
721
|
+
// The name is a valid key, so that the suggestion is valid.
|
|
722
|
+
const match = /^[\t ]*\[\[?(?<name>[A-Z_a-z][\w\-]*)\]\]?[\t ]*$/v.exec(line);
|
|
763
723
|
if (match === null) {
|
|
764
724
|
return;
|
|
765
725
|
}
|
|
766
726
|
const { name } = match.groups;
|
|
767
|
-
this.#fail(`There are no table headers. Write the table as an object, as in ${abbreviate(name)}: {…}
|
|
727
|
+
this.#fail(`There are no table headers. Write the table as an object, as in ${abbreviate(name)}: {…}`);
|
|
768
728
|
}
|
|
769
729
|
/*
|
|
770
730
|
A YAML literal block scalar indicator, as in `key: |` or `key: |-`, at the end of its line.
|
|
@@ -32,7 +32,7 @@ export type StringifyOptions = {
|
|
|
32
32
|
readonly canonical?: boolean;
|
|
33
33
|
};
|
|
34
34
|
/**
|
|
35
|
-
Serialize an object or an array to SOML. Nesting is written with braces and tabs, and the output ends with one line feed. Comments
|
|
35
|
+
Serialize an object or an array to SOML. Nesting is written with braces and tabs, and the output ends with one line feed. Comments and block strings are never written.
|
|
36
36
|
|
|
37
37
|
Members keep the order of `value`, which reads better in a file for people, and every other rule of canonical form is followed. With the `canonical: true` option, members are sorted by key, and the output is canonical form: two equal values produce the same bytes, so it can be hashed, signed, or compared.
|
|
38
38
|
|
|
@@ -27,7 +27,7 @@ const ESCAPES = new Map([
|
|
|
27
27
|
['\t', String.raw `\t`],
|
|
28
28
|
]);
|
|
29
29
|
/**
|
|
30
|
-
Serialize an object or an array to SOML. Nesting is written with braces and tabs, and the output ends with one line feed. Comments
|
|
30
|
+
Serialize an object or an array to SOML. Nesting is written with braces and tabs, and the output ends with one line feed. Comments and block strings are never written.
|
|
31
31
|
|
|
32
32
|
Members keep the order of `value`, which reads better in a file for people, and every other rule of canonical form is followed. With the `canonical: true` option, members are sorted by key, and the output is canonical form: two equal values produce the same bytes, so it can be hashed, signed, or compared.
|
|
33
33
|
|
package/distribution/tree.d.ts
CHANGED
|
@@ -38,7 +38,7 @@ type Located<Type extends string> = {
|
|
|
38
38
|
/**
|
|
39
39
|
A token. `value` is its source text.
|
|
40
40
|
|
|
41
|
-
A `Punctuator` is one of `{`, `}`, `[`, `]`, `:`,
|
|
41
|
+
A `Punctuator` is one of `{`, `}`, `[`, `]`, `:`, and `,`. A `Keyword` is `true`, `false`, or `null`. `infinity` and `-infinity` are `Float` tokens, and a quoted key is a `String` token.
|
|
42
42
|
*/
|
|
43
43
|
export type Token = Located<'Punctuator' | 'BareKey' | 'String' | 'Integer' | 'Float' | 'Keyword' | 'Instant' | 'Duration'> & {
|
|
44
44
|
readonly value: string;
|
|
@@ -84,7 +84,7 @@ A `key: value` member of an object.
|
|
|
84
84
|
*/
|
|
85
85
|
export type MemberNode = Located<'Member'> & {
|
|
86
86
|
/**
|
|
87
|
-
The key
|
|
87
|
+
The key.
|
|
88
88
|
*/
|
|
89
89
|
readonly key: KeyNode;
|
|
90
90
|
/**
|
|
@@ -102,24 +102,15 @@ export type ArrayNode = Located<'Array'> & {
|
|
|
102
102
|
readonly elements: readonly ValueNode[];
|
|
103
103
|
};
|
|
104
104
|
/**
|
|
105
|
-
A key
|
|
105
|
+
A key.
|
|
106
106
|
*/
|
|
107
107
|
export type KeyNode = Located<'Key'> & {
|
|
108
|
-
/**
|
|
109
|
-
The segments between the dots, in source order.
|
|
110
|
-
*/
|
|
111
|
-
readonly segments: readonly KeySegmentNode[];
|
|
112
|
-
};
|
|
113
|
-
/**
|
|
114
|
-
One segment of a key, which is the whole key unless it is dotted.
|
|
115
|
-
*/
|
|
116
|
-
export type KeySegmentNode = Located<'KeySegment'> & {
|
|
117
108
|
/**
|
|
118
109
|
The decoded key.
|
|
119
110
|
*/
|
|
120
111
|
readonly value: string;
|
|
121
112
|
/**
|
|
122
|
-
How the
|
|
113
|
+
How the key is written: bare, as `'...'`, or as `"..."`.
|
|
123
114
|
*/
|
|
124
115
|
readonly style: 'bare' | 'literal' | 'escaped';
|
|
125
116
|
};
|
|
@@ -221,7 +212,7 @@ export type ValueNode = ObjectNode | ArrayNode | StringNode | IntegerNode | Floa
|
|
|
221
212
|
/**
|
|
222
213
|
Any node in a tree.
|
|
223
214
|
*/
|
|
224
|
-
export type Node = DocumentNode | MemberNode | KeyNode |
|
|
215
|
+
export type Node = DocumentNode | MemberNode | KeyNode | ValueNode;
|
|
225
216
|
/**
|
|
226
217
|
The properties of each node type that hold its child nodes, in source order, for tools that walk the tree, such as an ESLint language plugin.
|
|
227
218
|
|
|
@@ -240,7 +231,7 @@ export declare const visitorKeys: Readonly<Record<Node['type'], readonly string[
|
|
|
240
231
|
/**
|
|
241
232
|
Parse a document into a syntax tree, for tools such as linters and formatters. Returns a `Document` node, which also holds every token and comment.
|
|
242
233
|
|
|
243
|
-
Every node, token, and comment has a `range`, which is `[start, end]` as UTF-16 offsets into `text`, and a `loc`, which is `{start: {line, column}, end: {line, column}}`, with a 1-based line and a 0-based column in UTF-16 code units, as in ESTree. A `ParseError` counts its column differently, for people to read, so use its `offset` to find the position in the tree. A scalar node or a key
|
|
234
|
+
Every node, token, and comment has a `range`, which is `[start, end]` as UTF-16 offsets into `text`, and a `loc`, which is `{start: {line, column}, end: {line, column}}`, with a 1-based line and a 0-based column in UTF-16 code units, as in ESTree. A `ParseError` counts its column differently, for people to read, so use its `offset` to find the position in the tree. A scalar node or a key shares its `range` and `loc` with its token, and other nodes, except `Document`, share the positions in `loc` with their first and last token, so treat them as read-only.
|
|
244
235
|
|
|
245
236
|
@param text - The document.
|
|
246
237
|
@returns The root node, which also holds every token and comment.
|
package/distribution/tree.js
CHANGED
|
@@ -1,5 +1,5 @@
|
|
|
1
1
|
import { parseWithTimes, } from "./parse.js";
|
|
2
|
-
import { isBareKeyCharacter, isSpace, findNumberEnd, findLineEnd, findBlockStringEnd, LF, DOUBLE_QUOTE, HASH, SINGLE_QUOTE, ASTERISK, COMMA,
|
|
2
|
+
import { isBareKeyCharacter, isSpace, findNumberEnd, findLineEnd, findBlockStringEnd, LF, DOUBLE_QUOTE, HASH, SINGLE_QUOTE, ASTERISK, COMMA, SLASH, BACKSLASH, OPEN_BRACKET, CLOSE_BRACKET, OPEN_BRACE, CLOSE_BRACE, } from "./shared.js";
|
|
3
3
|
/**
|
|
4
4
|
The properties of each node type that hold its child nodes, in source order, for tools that walk the tree, such as an ESLint language plugin.
|
|
5
5
|
|
|
@@ -20,8 +20,7 @@ export const visitorKeys = Object.freeze({
|
|
|
20
20
|
Object: Object.freeze(['members']),
|
|
21
21
|
Member: Object.freeze(['key', 'value']),
|
|
22
22
|
Array: Object.freeze(['elements']),
|
|
23
|
-
Key: Object.freeze([
|
|
24
|
-
KeySegment: Object.freeze([]),
|
|
23
|
+
Key: Object.freeze([]),
|
|
25
24
|
String: Object.freeze([]),
|
|
26
25
|
Integer: Object.freeze([]),
|
|
27
26
|
Float: Object.freeze([]),
|
|
@@ -40,7 +39,7 @@ const RADIX = new Map([['x', 16], ['o', 8], ['b', 2]]);
|
|
|
40
39
|
/**
|
|
41
40
|
Parse a document into a syntax tree, for tools such as linters and formatters. Returns a `Document` node, which also holds every token and comment.
|
|
42
41
|
|
|
43
|
-
Every node, token, and comment has a `range`, which is `[start, end]` as UTF-16 offsets into `text`, and a `loc`, which is `{start: {line, column}, end: {line, column}}`, with a 1-based line and a 0-based column in UTF-16 code units, as in ESTree. A `ParseError` counts its column differently, for people to read, so use its `offset` to find the position in the tree. A scalar node or a key
|
|
42
|
+
Every node, token, and comment has a `range`, which is `[start, end]` as UTF-16 offsets into `text`, and a `loc`, which is `{start: {line, column}, end: {line, column}}`, with a 1-based line and a 0-based column in UTF-16 code units, as in ESTree. A `ParseError` counts its column differently, for people to read, so use its `offset` to find the position in the tree. A scalar node or a key shares its `range` and `loc` with its token, and other nodes, except `Document`, share the positions in `loc` with their first and last token, so treat them as read-only.
|
|
44
43
|
|
|
45
44
|
@param text - The document.
|
|
46
45
|
@returns The root node, which also holds every token and comment.
|
|
@@ -75,7 +74,7 @@ function readScalar(raw) {
|
|
|
75
74
|
return parseWithTimes(`[${raw}]`)[0];
|
|
76
75
|
}
|
|
77
76
|
/*
|
|
78
|
-
Every node is an object literal with the same property order, `type`, its own fields, `range`, and `loc`, rather than one generic function that spreads the fields, which is several times slower. A node that is one token, such as a scalar or a key
|
|
77
|
+
Every node is an object literal with the same property order, `type`, its own fields, `range`, and `loc`, rather than one generic function that spreads the fields, which is several times slower. A node that is one token, such as a scalar or a key, shares the token's `range` and `loc`.
|
|
79
78
|
*/
|
|
80
79
|
class TreeBuilder {
|
|
81
80
|
#source;
|
|
@@ -236,27 +235,12 @@ class TreeBuilder {
|
|
|
236
235
|
};
|
|
237
236
|
}
|
|
238
237
|
#key() {
|
|
239
|
-
const segments = [this.#keySegment()];
|
|
240
|
-
while (this.#code() === DOT) {
|
|
241
|
-
this.#punctuator();
|
|
242
|
-
segments.push(this.#keySegment());
|
|
243
|
-
}
|
|
244
|
-
const first = segments[0];
|
|
245
|
-
const last = segments.at(-1);
|
|
246
|
-
return {
|
|
247
|
-
type: 'Key',
|
|
248
|
-
segments,
|
|
249
|
-
range: [first.range[0], last.range[1]],
|
|
250
|
-
loc: spanLocation(first, last),
|
|
251
|
-
};
|
|
252
|
-
}
|
|
253
|
-
#keySegment() {
|
|
254
238
|
const start = this.#index;
|
|
255
239
|
const code = this.#code();
|
|
256
240
|
if (code === SINGLE_QUOTE || code === DOUBLE_QUOTE) {
|
|
257
241
|
const { value, style, token: { range, loc } } = this.#singleLineString();
|
|
258
242
|
return {
|
|
259
|
-
type: '
|
|
243
|
+
type: 'Key',
|
|
260
244
|
value,
|
|
261
245
|
style,
|
|
262
246
|
range,
|
|
@@ -268,7 +252,7 @@ class TreeBuilder {
|
|
|
268
252
|
}
|
|
269
253
|
const { value, range, loc } = this.#token('BareKey', start, this.#index);
|
|
270
254
|
return {
|
|
271
|
-
type: '
|
|
255
|
+
type: 'Key',
|
|
272
256
|
value,
|
|
273
257
|
style: 'bare',
|
|
274
258
|
range,
|
package/package.json
CHANGED
package/readme.md
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# soml
|
|
2
2
|
|
|
3
|
-
> The reference parser, serializer
|
|
3
|
+
> The reference parser, serializer, and formatter for [SOML](https://soml.sh), a config format for humans
|
|
4
4
|
|
|
5
5
|
> [!NOTE]
|
|
6
6
|
> The format is a draft. See the [specification](https://github.com/soml-lang/soml/blob/main/spec.md).
|
|
@@ -40,7 +40,7 @@ replicas: 3
|
|
|
40
40
|
timeout: 30.0
|
|
41
41
|
grace: 1m30s
|
|
42
42
|
deployed-at: 2026-09-19T14:00:00Z
|
|
43
|
-
postgres
|
|
43
|
+
postgres: {host: 'db.internal'}
|
|
44
44
|
`);
|
|
45
45
|
//=> {
|
|
46
46
|
// name: 'api-gateway',
|
|
@@ -102,7 +102,7 @@ parse(`
|
|
|
102
102
|
name: 'api-gateway'
|
|
103
103
|
replicas: 3
|
|
104
104
|
timeout: 30.0
|
|
105
|
-
postgres
|
|
105
|
+
postgres: {host: 'db.internal'}
|
|
106
106
|
`);
|
|
107
107
|
//=> {name: 'api-gateway', replicas: 3n, timeout: 30, postgres: {host: 'db.internal'}}
|
|
108
108
|
```
|
|
@@ -150,7 +150,7 @@ tree.comments[0];
|
|
|
150
150
|
//=> {type: 'Line', value: ' The default', range: [11, 24], loc: {…}}
|
|
151
151
|
```
|
|
152
152
|
|
|
153
|
-
Every node, token, and comment has a `range`, which is `[start, end]` as UTF-16 offsets into `text`, and a `loc`, which is `{start: {line, column}, end: {line, column}}`, with a 1-based line and a 0-based column in UTF-16 code units, as in [ESTree](https://github.com/estree/estree). A [`ParseError`](#parseerror) counts its column differently, for people to read, so use its `offset` to find the position in the tree. A scalar node or a key
|
|
153
|
+
Every node, token, and comment has a `range`, which is `[start, end]` as UTF-16 offsets into `text`, and a `loc`, which is `{start: {line, column}, end: {line, column}}`, with a 1-based line and a 0-based column in UTF-16 code units, as in [ESTree](https://github.com/estree/estree). A [`ParseError`](#parseerror) counts its column differently, for people to read, so use its `offset` to find the position in the tree. A scalar node or a key shares its `range` and `loc` with its token, and other nodes, except `Document`, share the positions in `loc` with their first and last token, so treat them as read-only.
|
|
154
154
|
|
|
155
155
|
| Node | Fields |
|
|
156
156
|
|---|---|
|
|
@@ -158,8 +158,7 @@ Every node, token, and comment has a `range`, which is `[start, end]` as UTF-16
|
|
|
158
158
|
| `Object` | `members`; `braced`, which is `false` only for a top-level object without braces |
|
|
159
159
|
| `Member` | `key`; `value` |
|
|
160
160
|
| `Array` | `elements` |
|
|
161
|
-
| `Key` | `
|
|
162
|
-
| `KeySegment` | `value`, decoded; `style`: `'bare'`, `'literal'`, or `'escaped'` |
|
|
161
|
+
| `Key` | `value`, decoded; `style`: `'bare'`, `'literal'`, or `'escaped'` |
|
|
163
162
|
| `String` | `value`, decoded; `style`: `'literal'` or `'escaped'`; `block` |
|
|
164
163
|
| `Integer` | `value`, a `bigint`; `radix`: `2`, `8`, `10`, or `16` |
|
|
165
164
|
| `Float` | `value`, including `Infinity` and `-Infinity` |
|
|
@@ -168,7 +167,7 @@ Every node, token, and comment has a `range`, which is `[start, end]` as UTF-16
|
|
|
168
167
|
| `Instant` | `value`, a `Temporal.Instant`, made when it is first read, also by spreading or serializing the node |
|
|
169
168
|
| `Duration` | `value`, a `Temporal.Duration`, made when it is first read, also by spreading or serializing the node; `negative`, whether it is written with a `-`; `parts`, each `{number, unit}` as written, so `-1h1.5m` has the parts `{number: '1', unit: 'h'}` and `{number: '1.5', unit: 'm'}` |
|
|
170
169
|
|
|
171
|
-
A token is `{type, value, range, loc}`, where `type` is `'Punctuator'` (for `{`, `}`, `[`, `]`, `:`,
|
|
170
|
+
A token is `{type, value, range, loc}`, where `type` is `'Punctuator'` (for `{`, `}`, `[`, `]`, `:`, and `,`), `'BareKey'`, `'String'`, `'Integer'`, `'Float'`, `'Keyword'` (for `true`, `false`, and `null`), `'Instant'`, or `'Duration'`, and `value` is its source text. `infinity` and `-infinity` are `'Float'` tokens, and a quoted key is a `'String'` token. A comment is `{type, value, range, loc}`, where `type` is `'Line'` or `'Block'`, and `value` is the text without `#`, or without `/*` and `*/`.
|
|
172
171
|
|
|
173
172
|
### visitorKeys
|
|
174
173
|
|
|
@@ -188,7 +187,7 @@ visitorKeys.Integer;
|
|
|
188
187
|
|
|
189
188
|
Format a document. Returns the document with its layout normalized, ending with one line feed.
|
|
190
189
|
|
|
191
|
-
The layout follows the [formatter](https://github.com/soml-lang/soml/blob/main/spec.md#formatting) in the specification: one tab per level, every member and item on its own line with no commas, and no trailing whitespace or runs of blank lines. An object or an array whose brackets are on one line stays on one line, as in `ports: [80, 443]`, with a comma and a space between its members or items. To give a one-line container one member or item per line, put a line break anywhere inside it. Comments, member order,
|
|
190
|
+
The layout follows the [formatter](https://github.com/soml-lang/soml/blob/main/spec.md#formatting) in the specification: one tab per level, every member and item on its own line with no commas, and no trailing whitespace or runs of blank lines. An object or an array whose brackets are on one line stays on one line, as in `ports: [80, 443]`, with a comma and a space between its members or items. To give a one-line container one member or item per line, put a line break anywhere inside it. Comments, member order, block strings, and the spelling of every value stay as they are, so the value never changes. The one change inside a block comment is the same layout rule: trailing whitespace is removed, and runs of blank lines collapse to one.
|
|
192
191
|
|
|
193
192
|
Throws the same [`ParseError`](#parseerror) as `parse()`, and a `TypeError` when `text` is not a string.
|
|
194
193
|
|
|
@@ -237,20 +236,20 @@ import {edit} from 'soml-lang';
|
|
|
237
236
|
edit('name: \'api\' # The service\nport: 8080\n', ['port'], 9090n);
|
|
238
237
|
//=> "name: 'api' # The service\nport: 9090\n"
|
|
239
238
|
|
|
240
|
-
edit('postgres
|
|
241
|
-
//=> "postgres
|
|
239
|
+
edit('postgres: {host: \'db\'}\n', ['postgres', 'port'], 5432n);
|
|
240
|
+
//=> "postgres: {host: 'db', port: 5432}\n"
|
|
242
241
|
|
|
243
242
|
edit('a: 1\nb: 2\n', ['a'], undefined);
|
|
244
243
|
//=> 'b: 2\n'
|
|
245
244
|
```
|
|
246
245
|
|
|
247
246
|
- A new value is written as `stringify()` writes it, at the indentation of its line. So `0xFF` that is replaced by `255n` becomes `255`, and `8080` becomes `8080.0` unless you pass `8080n` or the `integers: 'number'` option. The members of a new object keep their order, unless you pass the `canonical: true` option. In a container that is on one line, it is written on one line too, so `[1, 2]` with a new item `{a: 3n}` becomes `[1, 2, {a: 3}]`.
|
|
248
|
-
- A new member goes after the last member of its object
|
|
247
|
+
- A new member goes after the last member of its object. Missing objects on the way are created.
|
|
249
248
|
- A new item can be added at the end of an array, with the index that is its length.
|
|
250
249
|
- A new member or item goes on its own line without a comma, after the comments that the one before it owns (see below), and the commas of the other members and items stay as they are. When something follows the member or item before it on the same line, such as the closing bracket of the one-line container `{a: 1}`, it goes on that line after a comma, as in `{a: 1, b: 2}`. In an empty `[]` or `{}`, it goes on a line of its own, unless that container is inside a container on one line, so `a: [1, []]` becomes `a: [1, [2]]`.
|
|
251
250
|
- A removed member or item takes the comments it owns, which are the ones that `format()` keeps with it: the comments after it on its line, also after its comma when nothing else follows there, and the block comments before it on its line, after the comma or bracket before it. So removing `2` from `[1, /* note */ 2]` gives `[1]`. A comment on a line of its own belongs to no member or item, so it stays. A removed member or item also removes its lines when nothing else is on them. A line that a block comment after it continues onto counts as one of its lines. Removing every item of a container closes it up to `[]` or `{}`, unless a comment is left inside.
|
|
252
251
|
- A removed member or item takes its comma with it. When the removed items are the last ones and have no comma after them, they take the comma directly before them instead, when only spaces, tabs, and the comments they own are between, so `[1, 2]` becomes `[1]`.
|
|
253
|
-
- Removing the
|
|
252
|
+
- Removing the only member of a document without braces leaves `{}`.
|
|
254
253
|
|
|
255
254
|
Comments inside a value that is replaced or removed are removed with it. In a layout that `format()` never writes, an edit can leave odd spacing, such as a new item after a closing block string delimiter on its line that is indented differently from its neighbors. The document is always valid and has the right value. Removing a value that does not exist changes nothing and is not an error: a missing member, a missing object or array on the way, or an index at or past the end of its array. So removing the same path twice is safe, and `edit(text, path, undefined) === text` tells whether something was removed.
|
|
256
255
|
|
|
@@ -268,7 +267,7 @@ The document.
|
|
|
268
267
|
|
|
269
268
|
Type: `Array<string | number>`
|
|
270
269
|
|
|
271
|
-
The keys and array indexes that lead to the value, such as `['servers', 0, 'port']`.
|
|
270
|
+
The keys and array indexes that lead to the value, such as `['servers', 0, 'port']`.
|
|
272
271
|
|
|
273
272
|
#### value
|
|
274
273
|
|
|
@@ -306,7 +305,7 @@ edit('a: 1\n', ['b'], {y: 1n, x: 2n}, {canonical: true});
|
|
|
306
305
|
|
|
307
306
|
### stringify(value, options?)
|
|
308
307
|
|
|
309
|
-
Serialize an object or an array to SOML. Nesting is written with braces and tabs, and the output ends with one line feed. Comments
|
|
308
|
+
Serialize an object or an array to SOML. Nesting is written with braces and tabs, and the output ends with one line feed. Comments and block strings are never written.
|
|
310
309
|
|
|
311
310
|
Members keep the order of `value`, which reads better in a file for people, and every other rule of canonical form is followed. With the [`canonical: true`](#canonical-1) option, members are sorted by key, and the output is canonical form: two equal values produce the same bytes, so it can be hashed, signed, or compared.
|
|
312
311
|
|
|
@@ -477,7 +476,7 @@ Up to three lines ending at the error, with a caret under the position. A long l
|
|
|
477
476
|
|
|
478
477
|
## Limits
|
|
479
478
|
|
|
480
|
-
- **Nesting is limited to 100 levels**, as the spec requires: every document up to 100 levels is accepted, by `parse()`, `parseTree()`, `format()`, and `edit()`, and every deeper one is rejected. `stringify()` refuses a deeper value, and `edit()` refuses a change that would make one. The document's own collection is level 1,
|
|
479
|
+
- **Nesting is limited to 100 levels**, as the spec requires: every document up to 100 levels is accepted, by `parse()`, `parseTree()`, `format()`, and `edit()`, and every deeper one is rejected. `stringify()` refuses a deeper value, and `edit()` refuses a change that would make one. The document's own collection is level 1, so `a: {b: {c: 1}}` has a depth of 3.
|
|
481
480
|
|
|
482
481
|
## Conformance suite
|
|
483
482
|
|
|
@@ -504,7 +503,7 @@ The same 2.5 MB of config-shaped data on an Apple M-series machine with Node.js
|
|
|
504
503
|
| smol-toml | 100 MB/s | 135 MB/s |
|
|
505
504
|
| json5 | 22 MB/s | 85 MB/s |
|
|
506
505
|
|
|
507
|
-
A hand-written style document (comments,
|
|
506
|
+
A hand-written style document (comments, block strings, hex ints, instants) parses at about 110 MB/s, and minified one-line input at about 120 MB/s. The `stringify` figure is with `canonical: true`, which sorts every object's keys, a step smol-toml does not need. The default keeps member order and skips that step, so it is faster.
|
|
508
507
|
|
|
509
508
|
`parseTree()` runs at about 35 MB/s, and `format()` at about 20 to 25 MB/s, also on an already formatted document. Both are slower than `parse()`, because the tree has a node and a location for every token.
|
|
510
509
|
|