tree-sitter-ktav 0.3.0 → 0.6.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +2 -1
- package/README.ru.md +1 -1
- package/README.zh.md +1 -1
- package/grammar.js +231 -137
- package/package.json +2 -2
- package/queries/highlights.scm +13 -23
- package/src/grammar.json +415 -147
- package/src/node-types.json +224 -33
- package/src/parser.c +4272 -2168
package/README.md
CHANGED
|
@@ -171,7 +171,8 @@ rejects (mostly missing-whitespace-after-marker — § 6.10).
|
|
|
171
171
|
|
|
172
172
|
## License
|
|
173
173
|
|
|
174
|
-
MIT — see
|
|
174
|
+
Dual-licensed under **MIT OR Apache-2.0** — see
|
|
175
|
+
[LICENSE-MIT](LICENSE-MIT) and [LICENSE-APACHE](LICENSE-APACHE).
|
|
175
176
|
|
|
176
177
|
## Other Ktav repositories
|
|
177
178
|
|
package/README.ru.md
CHANGED
package/README.zh.md
CHANGED
package/grammar.js
CHANGED
|
@@ -1,24 +1,43 @@
|
|
|
1
1
|
/**
|
|
2
2
|
* Tree-sitter grammar for Ktav (כְּתָב) — the Written Configuration Format.
|
|
3
3
|
*
|
|
4
|
-
* Spec: https://github.com/ktav-lang/spec/blob/main/versions/0.
|
|
4
|
+
* Spec: https://github.com/ktav-lang/spec/blob/main/versions/0.6/spec.md
|
|
5
5
|
*
|
|
6
6
|
* Ktav is line-oriented. Every line is one of:
|
|
7
7
|
* - blank
|
|
8
|
-
* - a comment (
|
|
9
|
-
* - a key:value pair (with markers
|
|
8
|
+
* - a comment (`## ...`) ← NOTE: double-hash in 0.5.0; single `#` is content
|
|
9
|
+
* - a key:value pair (with markers `:` or `::`)
|
|
10
10
|
* - a structural opener / closer for compounds (`{`, `}`, `[`, `]`,
|
|
11
11
|
* `(`, `((`, `)`, `))`)
|
|
12
|
-
* - an array item (inside an open `[` array)
|
|
12
|
+
* - an array item (inside an open `[` array, or at the top level)
|
|
13
13
|
* - raw content of a multi-line string
|
|
14
14
|
*
|
|
15
|
+
* Changes from 0.5.0 to 0.6.0:
|
|
16
|
+
* - Keys now process escape sequences (§ 3.7). The escape table grows
|
|
17
|
+
* to 10 entries by adding `\.` (literal dot — does NOT split the
|
|
18
|
+
* dotted path) and `\:` (literal colon — does NOT act as the
|
|
19
|
+
* key/value separator). `\` becomes the escape lead in keys.
|
|
20
|
+
* Breaking: a raw `\` in a key now requires `\\`.
|
|
21
|
+
*
|
|
22
|
+
* Changes from 0.3.0 (spec 0.1.1) to 0.5.0:
|
|
23
|
+
* - Comment marker changed from `#` to `##`. Single `#` is now a
|
|
24
|
+
* content byte (allowed in keys and scalar values).
|
|
25
|
+
* - Typed markers `:i` and `:f` removed. Only `:` and `::` remain.
|
|
26
|
+
* - Inline compounds: `{key: value, ...}` and `[v1, v2, ...]` are
|
|
27
|
+
* now valid as pair values or array items (inline_object /
|
|
28
|
+
* inline_array rules).
|
|
29
|
+
* - Number literals: hex (`0x`), octal (`0o`), binary (`0b`), decimal
|
|
30
|
+
* with underscore separators, and floats with `.` or exponent.
|
|
31
|
+
* These are captured as distinct node kinds for highlighting.
|
|
32
|
+
* - Escape sequences inside inline scalars: `\\`, `\,`, `\}`, `\]`,
|
|
33
|
+
* `\{`, `\[`, `\n`, `\r`.
|
|
34
|
+
*
|
|
15
35
|
* Strategy:
|
|
16
36
|
* - Newlines are explicit (`_newline`) and structural openers/closers
|
|
17
37
|
* are tokens that include the trailing whitespace + newline so they
|
|
18
38
|
* can NEVER be confused with a scalar starting with the same byte.
|
|
19
|
-
* - The
|
|
20
|
-
*
|
|
21
|
-
* > `:`).
|
|
39
|
+
* - The two pair separators (`:`, `::`) are recognized by the lexer
|
|
40
|
+
* with longest-match precedence (`::` > `:`).
|
|
22
41
|
* - The mandatory-whitespace-after-marker rule (§ 6.10) is enforced
|
|
23
42
|
* by an external scanner token `_marker_ws`, which is a zero-width
|
|
24
43
|
* assertion that only succeeds when the byte right after the
|
|
@@ -31,6 +50,8 @@
|
|
|
31
50
|
* a parse error.
|
|
32
51
|
* - Multi-line string content is captured as a sequence of opaque
|
|
33
52
|
* "raw lines" up to the matching terminator.
|
|
53
|
+
* - Inline compound values are fully parsed; escape sequences inside
|
|
54
|
+
* inline scalars are captured as `escape_sequence` nodes.
|
|
34
55
|
* - Indentation is not significant (matches the spec).
|
|
35
56
|
*/
|
|
36
57
|
|
|
@@ -56,29 +77,12 @@ module.exports = grammar({
|
|
|
56
77
|
|
|
57
78
|
rules: {
|
|
58
79
|
// The top-level document is a sequence of lines. Per spec § 5.0.1
|
|
59
|
-
//
|
|
60
|
-
//
|
|
61
|
-
//
|
|
62
|
-
// parser of trees: rather than encode the spec's "first content
|
|
63
|
-
// line decides" rule structurally, we accept either kind of line
|
|
64
|
-
// anywhere at the top level. The reference parser performs the
|
|
65
|
-
// semantic § 5.0.1 dispatch and rejects mixed roots; the grammar's
|
|
66
|
-
// job here is only to produce a clean syntax tree for both shapes
|
|
67
|
-
// (and for editors / LSPs) without spurious ERROR nodes on bare
|
|
68
|
-
// top-level scalars or `:i`/`:f`/`::` lines.
|
|
80
|
+
// the root may be either an Object or an Array. Tree-sitter accepts
|
|
81
|
+
// both kinds of line anywhere; semantic dispatch is left to the
|
|
82
|
+
// reference parser.
|
|
69
83
|
source_file: $ => repeat($._line),
|
|
70
84
|
|
|
71
85
|
// ---- Top-level lines ----
|
|
72
|
-
//
|
|
73
|
-
// `top_array_item` covers the top-level Array case (§ 5.0.1): a
|
|
74
|
-
// bare scalar, a typed-marker item (`:: …`, `:i …`, `:f …`), a
|
|
75
|
-
// lone `{` / `[` opener, or a `(` / `((` multi-line opener. It
|
|
76
|
-
// is structurally identical to `array_item` (used inside `[ ]`)
|
|
77
|
-
// but is kept as a distinct rule so we can apply a lexer-level
|
|
78
|
-
// disambiguation: a line that classifies as a pair (§ 5.3) must
|
|
79
|
-
// remain a pair at the top level (spec § 5.0.1 step 2), so
|
|
80
|
-
// `top_array_item`'s plain bare-scalar branch is given lower
|
|
81
|
-
// precedence than `object_pair` via `prec.dynamic`.
|
|
82
86
|
_line: $ => choice(
|
|
83
87
|
$.comment,
|
|
84
88
|
$.blank_line,
|
|
@@ -92,27 +96,25 @@ module.exports = grammar({
|
|
|
92
96
|
|
|
93
97
|
// ---- Comment ----
|
|
94
98
|
//
|
|
95
|
-
//
|
|
96
|
-
//
|
|
97
|
-
//
|
|
98
|
-
//
|
|
99
|
-
|
|
100
|
-
comment: $ => token(prec(1, /#[^\r\n]*\r?\n/)),
|
|
99
|
+
// In spec 0.5.0 a comment starts with `##` (two hashes). A single
|
|
100
|
+
// `#` is ordinary content. The token captures the whole line
|
|
101
|
+
// including the trailing newline to beat `_top_scalar_text` at the
|
|
102
|
+
// lexer's longest-match step.
|
|
103
|
+
comment: $ => token(prec(1, /##[^\r\n]*\r?\n/)),
|
|
101
104
|
|
|
102
105
|
// ---- Object pair ----
|
|
103
106
|
//
|
|
104
107
|
// After the separator, the external `_marker_ws` token asserts
|
|
105
|
-
// that the next byte is whitespace, CR, LF, or EOF
|
|
106
|
-
// § 6.10 (mandatory whitespace after marker): `key:value` fails.
|
|
108
|
+
// that the next byte is whitespace, CR, LF, or EOF (§ 6.10).
|
|
107
109
|
object_pair: $ => choice(
|
|
108
|
-
// After
|
|
109
|
-
//
|
|
110
|
-
//
|
|
110
|
+
// After `::` the body is a literal scalar (NOT dispatched through
|
|
111
|
+
// compound-opener / multi-line dispatch). `raw_scalar` accepts any
|
|
112
|
+
// line content including `(`, `{`, `[`-starting text.
|
|
111
113
|
seq(
|
|
112
114
|
field('key', $.key),
|
|
113
|
-
field('separator',
|
|
115
|
+
field('separator', $.sep_raw),
|
|
114
116
|
$._marker_ws,
|
|
115
|
-
field('value', choice($.empty_value, $.
|
|
117
|
+
field('value', choice($.empty_value, $.raw_scalar)),
|
|
116
118
|
),
|
|
117
119
|
// After `:` the body goes through the full § 5.2 dispatch.
|
|
118
120
|
seq(
|
|
@@ -126,27 +128,39 @@ module.exports = grammar({
|
|
|
126
128
|
),
|
|
127
129
|
),
|
|
128
130
|
|
|
129
|
-
//
|
|
130
|
-
// over `:`, `::` is given higher precedence; same for `:i`/`:f`.
|
|
131
|
+
// `::` wins over `:` via higher precedence.
|
|
131
132
|
sep_raw: $ => token(prec(3, '::')),
|
|
132
|
-
sep_int: $ => token(prec(2, ':i')),
|
|
133
|
-
sep_float: $ => token(prec(2, ':f')),
|
|
134
133
|
sep_string: $ => token(prec(1, ':')),
|
|
135
134
|
|
|
136
135
|
// ---- Keys ----
|
|
137
136
|
key: $ => choice(
|
|
138
|
-
$.
|
|
137
|
+
$._spaced_key,
|
|
139
138
|
$.dotted_key,
|
|
140
139
|
),
|
|
141
140
|
|
|
141
|
+
// A key may contain internal whitespace (spec 0.5.0 § 4): the run of
|
|
142
|
+
// space-separated segments up to the separator is one key
|
|
143
|
+
// (`multi word key: value`). The inter-segment spaces are `extras`,
|
|
144
|
+
// so the `key` node still spans the whole text with no named children
|
|
145
|
+
// (renders as `(key)`, same as a single-segment key).
|
|
146
|
+
_spaced_key: $ => prec.left(repeat1($._key_segment)),
|
|
147
|
+
|
|
142
148
|
dotted_key: $ => prec.left(seq(
|
|
143
149
|
$._key_segment,
|
|
144
150
|
repeat1(seq('.', $._key_segment)),
|
|
145
151
|
)),
|
|
146
152
|
|
|
147
|
-
// Key segment:
|
|
148
|
-
//
|
|
149
|
-
|
|
153
|
+
// Key segment (spec 0.6.0 § 4): a non-empty run of plain key bytes
|
|
154
|
+
// and/or escape sequences. Plain key bytes exclude whitespace, the
|
|
155
|
+
// bracket/paren/brace bytes, `:`, `,`, the dotted-path separator `.`,
|
|
156
|
+
// `#` (reserved for comment marker `##`), and the escape lead `\`.
|
|
157
|
+
// Those structural bytes — including `\.` and `\:` — can now appear
|
|
158
|
+
// inside a key when escaped (§ 3.7, expanded in 0.6.0 to include
|
|
159
|
+
// `\.` and `\:`).
|
|
160
|
+
//
|
|
161
|
+
// The dotted-path separator is an UNescaped `.`; an unescaped `\`
|
|
162
|
+
// is always the start of an escape sequence (10 forms in 0.6.0).
|
|
163
|
+
_key_segment: $ => /([^\s\[\]\{\}\(\):#,.\r\n\\]|\\[\\,\}\]\{\[nr.:])+/,
|
|
150
164
|
|
|
151
165
|
// ---- Value line ----
|
|
152
166
|
_value_line: $ => choice(
|
|
@@ -160,8 +174,14 @@ module.exports = grammar({
|
|
|
160
174
|
$.empty_array,
|
|
161
175
|
$.empty_paren,
|
|
162
176
|
$.empty_double_paren,
|
|
177
|
+
// Inline compounds (new in 0.5.0).
|
|
178
|
+
$.inline_object,
|
|
179
|
+
$.inline_array,
|
|
163
180
|
// Keywords (single token followed by newline).
|
|
164
181
|
$.keyword,
|
|
182
|
+
// Number literals (distinct nodes for highlighting).
|
|
183
|
+
$.integer,
|
|
184
|
+
$.float,
|
|
165
185
|
// Scalar — catch-all line content.
|
|
166
186
|
$.scalar,
|
|
167
187
|
),
|
|
@@ -176,25 +196,12 @@ module.exports = grammar({
|
|
|
176
196
|
empty_double_paren: $ => seq(token(prec(5, '(())')), $._newline),
|
|
177
197
|
|
|
178
198
|
// ---- Multi-line compounds ----
|
|
179
|
-
//
|
|
180
|
-
// Openers are tokens that include the rest of the line up to and
|
|
181
|
-
// including the newline, so they cannot be confused with a scalar
|
|
182
|
-
// starting with the same byte. Closers, by contrast, are split
|
|
183
|
-
// into the bracket character(s) plus the external `_strict_eol`
|
|
184
|
-
// token; this lets the scanner reject pathological lines like
|
|
185
|
-
// `} trailing_text\n` (§ 5.6.1 and the cleanliness rule for
|
|
186
|
-
// object/array closers).
|
|
187
199
|
open_brace: $ => token(prec(4, /\{[ \t]*\r?\n/)),
|
|
188
200
|
close_brace: $ => seq(token(prec(4, '}')), $._strict_eol),
|
|
189
201
|
open_bracket: $ => token(prec(4, /\[[ \t]*\r?\n/)),
|
|
190
202
|
close_bracket: $ => seq(token(prec(4, ']')), $._strict_eol),
|
|
191
203
|
open_paren: $ => token(prec(4, /\([ \t]*\r?\n/)),
|
|
192
204
|
open_dparen: $ => token(prec(5, /\(\([ \t]*\r?\n/)),
|
|
193
|
-
// close_paren / close_dparen are emitted by the external scanner as
|
|
194
|
-
// `_stripped_close` / `_verbatim_close`. They are context-sensitive:
|
|
195
|
-
// inside `(...)` only `)` closes (a `))` line is content); inside
|
|
196
|
-
// `((...))` only `))` closes (a single `)` line is content). The
|
|
197
|
-
// scanner consults `valid_symbols` to decide which form to attempt.
|
|
198
205
|
close_paren: $ => $._stripped_close,
|
|
199
206
|
close_dparen: $ => $._verbatim_close,
|
|
200
207
|
|
|
@@ -210,62 +217,128 @@ module.exports = grammar({
|
|
|
210
217
|
$.close_bracket,
|
|
211
218
|
),
|
|
212
219
|
|
|
213
|
-
// ----
|
|
220
|
+
// ---- Inline compounds (new in spec 0.5.0) ----
|
|
214
221
|
//
|
|
215
|
-
//
|
|
216
|
-
//
|
|
222
|
+
// `{key: value, key2: value2}` and `[v1, v2, v3]` are valid as a
|
|
223
|
+
// value on the right-hand side of a pair or as an array item.
|
|
224
|
+
// Trailing comma is allowed. Nesting is supported.
|
|
225
|
+
//
|
|
226
|
+
// These are followed by a newline (they consume the rest of the line).
|
|
227
|
+
inline_object: $ => seq(
|
|
228
|
+
'{',
|
|
229
|
+
optional($._inline_pair_list),
|
|
230
|
+
'}',
|
|
231
|
+
$._newline,
|
|
232
|
+
),
|
|
233
|
+
|
|
234
|
+
inline_array: $ => seq(
|
|
235
|
+
'[',
|
|
236
|
+
optional($._inline_item_list),
|
|
237
|
+
']',
|
|
238
|
+
$._newline,
|
|
239
|
+
),
|
|
240
|
+
|
|
241
|
+
_inline_pair_list: $ => seq(
|
|
242
|
+
$.inline_pair,
|
|
243
|
+
repeat(seq(',', $.inline_pair)),
|
|
244
|
+
optional(','),
|
|
245
|
+
),
|
|
246
|
+
|
|
247
|
+
// The value is optional: `{x:, y: 1}` and `{empty:}` are valid —
|
|
248
|
+
// a separator immediately followed by `,` or `}` is an empty value
|
|
249
|
+
// (spec 0.5.0 § 5.8).
|
|
250
|
+
inline_pair: $ => seq(
|
|
251
|
+
field('key', $.key),
|
|
252
|
+
field('separator', choice($.sep_raw, $.sep_string)),
|
|
253
|
+
optional(field('value', $.inline_value)),
|
|
254
|
+
),
|
|
255
|
+
|
|
256
|
+
_inline_item_list: $ => seq(
|
|
257
|
+
$.inline_value,
|
|
258
|
+
repeat(seq(',', $.inline_value)),
|
|
259
|
+
optional(','),
|
|
260
|
+
),
|
|
261
|
+
|
|
262
|
+
// An inline value is either a nested inline compound, or an inline
|
|
263
|
+
// scalar (which may contain escape sequences).
|
|
264
|
+
inline_value: $ => choice(
|
|
265
|
+
$.nested_inline_object,
|
|
266
|
+
$.nested_inline_array,
|
|
267
|
+
$.inline_scalar,
|
|
268
|
+
),
|
|
269
|
+
|
|
270
|
+
nested_inline_object: $ => seq(
|
|
271
|
+
'{',
|
|
272
|
+
optional($._inline_pair_list),
|
|
273
|
+
'}',
|
|
274
|
+
),
|
|
275
|
+
|
|
276
|
+
nested_inline_array: $ => seq(
|
|
277
|
+
'[',
|
|
278
|
+
optional($._inline_item_list),
|
|
279
|
+
']',
|
|
280
|
+
),
|
|
281
|
+
|
|
282
|
+
// An inline scalar is terminated by an unescaped `,`, `}`, or `]`.
|
|
283
|
+
// It may contain escape sequences (§ 3.7).
|
|
284
|
+
//
|
|
285
|
+
// The first chunk must NOT begin with `{` or `[`: a value position
|
|
286
|
+
// that opens with `{`/`[` is a nested compound, not a scalar. After
|
|
287
|
+
// the first character those bytes are ordinary literal content
|
|
288
|
+
// (`hello{world`, `mid[bracket`). This head/rest split keeps
|
|
289
|
+
// `nested_inline_*` vs `inline_scalar` unambiguous without sacrificing
|
|
290
|
+
// mid-value literal braces (§ 5.8).
|
|
291
|
+
inline_scalar: $ => seq(
|
|
292
|
+
choice($.escape_sequence, $._inline_scalar_head),
|
|
293
|
+
repeat(choice($.escape_sequence, $._inline_scalar_text)),
|
|
294
|
+
),
|
|
295
|
+
|
|
296
|
+
// Escape sequences recognised inside inline scalars (§ 3.7).
|
|
297
|
+
// In 0.6.0 the table grows to 10 forms by adding `\.` and `\:` —
|
|
298
|
+
// they yield literal `.` and `:` respectively. Inside a value these
|
|
299
|
+
// are redundant (a value's `.`/`:` is already a literal byte) but
|
|
300
|
+
// remain accepted for symmetry with key parsing.
|
|
301
|
+
escape_sequence: $ => token(choice(
|
|
302
|
+
'\\\\',
|
|
303
|
+
'\\,',
|
|
304
|
+
'\\}',
|
|
305
|
+
'\\]',
|
|
306
|
+
'\\{',
|
|
307
|
+
'\\[',
|
|
308
|
+
'\\n',
|
|
309
|
+
'\\r',
|
|
310
|
+
'\\.',
|
|
311
|
+
'\\:',
|
|
312
|
+
)),
|
|
313
|
+
|
|
314
|
+
// Leading chunk of a scalar: first byte excludes whitespace and the
|
|
315
|
+
// openers `{`/`[` (so a value starting with an opener is a nested
|
|
316
|
+
// compound), plus the usual `\` `,` `}` `]` / CR / LF. Subsequent
|
|
317
|
+
// bytes allow `{`/`[` as literal content.
|
|
318
|
+
_inline_scalar_head: $ => token(/[^\s\\,\{\[\}\]\r\n][^\\,\}\]\r\n]*/),
|
|
319
|
+
|
|
320
|
+
// Continuation text after the head (or after an escape): any byte
|
|
321
|
+
// except `\`, `,`, the closers `}` `]`, and CR / LF. Open delimiters
|
|
322
|
+
// `{`/`[` are allowed here as literal content.
|
|
323
|
+
_inline_scalar_text: $ => token(/[^\\,\}\]\r\n]+/),
|
|
324
|
+
|
|
325
|
+
// ---- Array items ----
|
|
217
326
|
array_item: $ => choice(
|
|
218
327
|
seq(
|
|
219
328
|
field('marker', $.sep_raw),
|
|
220
329
|
$._marker_ws,
|
|
221
|
-
field('value', choice($.empty_value, $.
|
|
222
|
-
),
|
|
223
|
-
seq(
|
|
224
|
-
field('marker', $.sep_int),
|
|
225
|
-
$._marker_ws,
|
|
226
|
-
field('value', choice($.empty_value, $.scalar)),
|
|
227
|
-
),
|
|
228
|
-
seq(
|
|
229
|
-
field('marker', $.sep_float),
|
|
230
|
-
$._marker_ws,
|
|
231
|
-
field('value', choice($.empty_value, $.scalar)),
|
|
330
|
+
field('value', choice($.empty_value, $.raw_scalar)),
|
|
232
331
|
),
|
|
233
332
|
// Plain value item — same set as object pair value.
|
|
234
333
|
field('value', $._value_line),
|
|
235
334
|
),
|
|
236
335
|
|
|
237
|
-
// Top-level array item (§ 5.0.1
|
|
238
|
-
//
|
|
239
|
-
// Structurally a sibling of `array_item` (used inside `[ ]`),
|
|
240
|
-
// but with two differences that realise spec § 5.0.1 step 2:
|
|
241
|
-
//
|
|
242
|
-
// * The plain bare-scalar branch uses `_top_scalar_text`, a
|
|
243
|
-
// token that disallows `:` anywhere in the line. This makes
|
|
244
|
-
// any line containing a `:` (i.e. any pair-shaped line)
|
|
245
|
-
// parseable ONLY as `object_pair`, never as a top-level
|
|
246
|
-
// bare-scalar — which is what § 5.0.1 step 2 requires.
|
|
247
|
-
// Per spec note: to force a colon-bearing scalar at the root
|
|
248
|
-
// to be an Array item, use the raw marker (`:: foo: bar`).
|
|
249
|
-
//
|
|
250
|
-
// * Compound openers (`{`, `[`, `(`, `((`) and the empty-inline
|
|
251
|
-
// compound forms (`{}` / `[]` / `()` / `(())`) and the
|
|
252
|
-
// keywords are still allowed unchanged — they are
|
|
253
|
-
// unambiguous lines that cannot be confused with a pair.
|
|
336
|
+
// Top-level array item (§ 5.0.1).
|
|
254
337
|
top_array_item: $ => choice(
|
|
255
338
|
seq(
|
|
256
339
|
field('marker', $.sep_raw),
|
|
257
340
|
$._marker_ws,
|
|
258
|
-
field('value', choice($.empty_value, $.
|
|
259
|
-
),
|
|
260
|
-
seq(
|
|
261
|
-
field('marker', $.sep_int),
|
|
262
|
-
$._marker_ws,
|
|
263
|
-
field('value', choice($.empty_value, $.scalar)),
|
|
264
|
-
),
|
|
265
|
-
seq(
|
|
266
|
-
field('marker', $.sep_float),
|
|
267
|
-
$._marker_ws,
|
|
268
|
-
field('value', choice($.empty_value, $.scalar)),
|
|
341
|
+
field('value', choice($.empty_value, $.raw_scalar)),
|
|
269
342
|
),
|
|
270
343
|
field('value', $.compound_object),
|
|
271
344
|
field('value', $.compound_array),
|
|
@@ -275,35 +348,18 @@ module.exports = grammar({
|
|
|
275
348
|
field('value', $.empty_array),
|
|
276
349
|
field('value', $.empty_paren),
|
|
277
350
|
field('value', $.empty_double_paren),
|
|
351
|
+
field('value', $.inline_object),
|
|
352
|
+
field('value', $.inline_array),
|
|
278
353
|
field('value', $.keyword),
|
|
279
|
-
|
|
280
|
-
|
|
281
|
-
// (no `:` allowed). Wrapping in `top_scalar` keeps the AST
|
|
282
|
-
// tidy: tools see exactly the same `(scalar)` shape as
|
|
283
|
-
// elsewhere, with the only difference being which lexer
|
|
284
|
-
// rule produced the inner token.
|
|
354
|
+
field('value', $.integer),
|
|
355
|
+
field('value', $.float),
|
|
285
356
|
field('value', $.top_scalar),
|
|
286
357
|
),
|
|
287
358
|
|
|
288
|
-
// `top_scalar`
|
|
289
|
-
//
|
|
290
|
-
// case. The token swallows the trailing newline, which makes
|
|
291
|
-
// it strictly longer than the `_key_segment` token a pair
|
|
292
|
-
// would emit (`_key_segment` stops at the first `:` / `\n` /
|
|
293
|
-
// structural byte). Two consequences:
|
|
294
|
-
//
|
|
295
|
-
// 1. On a colon-free line such as `foo\n`, the whole-line
|
|
296
|
-
// `_top_scalar_text` token (length 4) wins the lexer's
|
|
297
|
-
// longest-match against `_key_segment` (length 3) — so
|
|
298
|
-
// the parser commits to the top-level Array-item path,
|
|
299
|
-
// not the always-ERROR pair-without-separator path.
|
|
300
|
-
// 2. On a pair-shaped line such as `name: Russia\n`, the
|
|
301
|
-
// `_top_scalar_text` regex CANNOT match (it forbids
|
|
302
|
-
// `:`), so `_key_segment` is the only viable token at
|
|
303
|
-
// line-start and the parser correctly enters
|
|
304
|
-
// `object_pair` (spec § 5.0.1 step 2).
|
|
359
|
+
// `top_scalar` — bare-scalar at the document root. Forbids `:` so
|
|
360
|
+
// that pair-shaped lines always parse as `object_pair`.
|
|
305
361
|
top_scalar: $ => $._top_scalar_text,
|
|
306
|
-
_top_scalar_text: $ => token(/[^\s:\r\n][^:\r\n]*\r?\n/),
|
|
362
|
+
_top_scalar_text: $ => token(/[^\s:\{\[\(\r\n][^:\r\n]*\r?\n/),
|
|
307
363
|
|
|
308
364
|
// ---- Multi-line strings ----
|
|
309
365
|
multiline_stripped: $ => seq(
|
|
@@ -318,11 +374,6 @@ module.exports = grammar({
|
|
|
318
374
|
$.close_dparen,
|
|
319
375
|
),
|
|
320
376
|
|
|
321
|
-
// A content line inside a multi-line string. Captured as one
|
|
322
|
-
// token. The closer tokens (`)`, `))` plus _strict_eol) win at
|
|
323
|
-
// the LR(1) boundary because they require their own line; any
|
|
324
|
-
// line whose content includes more than just the terminator
|
|
325
|
-
// falls through to this rule.
|
|
326
377
|
multiline_content_line: $ => token(prec(-1, /[^\r\n]*\r?\n/)),
|
|
327
378
|
|
|
328
379
|
// ---- Scalar (default value body, until end of line) ----
|
|
@@ -331,12 +382,55 @@ module.exports = grammar({
|
|
|
331
382
|
$._newline,
|
|
332
383
|
),
|
|
333
384
|
|
|
385
|
+
// `raw_scalar` is used exclusively after `::` (raw marker). It accepts
|
|
386
|
+
// any non-empty line content, including `(`, `{`, `[`-starting text.
|
|
387
|
+
// Per spec § 5.2: after `::` the body is NEVER dispatched as a
|
|
388
|
+
// compound opener or multi-line string opener — it is always a literal
|
|
389
|
+
// String value. This is a separate rule (not `scalar`) because
|
|
390
|
+
// `scalar`'s underlying `_scalar_text` deliberately excludes those
|
|
391
|
+
// opening bytes to avoid lexer ambiguity in the non-raw value context.
|
|
392
|
+
raw_scalar: $ => seq(
|
|
393
|
+
$._raw_scalar_text,
|
|
394
|
+
$._newline,
|
|
395
|
+
),
|
|
396
|
+
_raw_scalar_text: $ => /[^\s\r\n][^\r\n]*/,
|
|
397
|
+
|
|
334
398
|
// Scalar text: any non-whitespace, non-newline content up to end
|
|
335
|
-
// of line. `#`
|
|
336
|
-
//
|
|
337
|
-
//
|
|
338
|
-
//
|
|
339
|
-
|
|
399
|
+
// of line. Both `#` and `##` are allowed as content bytes in 0.5.0.
|
|
400
|
+
// Lines starting with `{` or `[` are always handled by structural
|
|
401
|
+
// rules (compound_object, compound_array, empty_object, empty_array,
|
|
402
|
+
// inline_object, inline_array), so `_scalar_text` explicitly excludes
|
|
403
|
+
// those opening bytes at position 0 to avoid the greedy-token
|
|
404
|
+
// ambiguity. Lines starting with `(` are handled by multiline or
|
|
405
|
+
// empty-paren rules likewise.
|
|
406
|
+
_scalar_text: $ => /[^\s\{\[\(\r\n][^\r\n]*/,
|
|
407
|
+
|
|
408
|
+
// ---- Number literals ----
|
|
409
|
+
//
|
|
410
|
+
// Captured as distinct node kinds so syntax highlighters can colour
|
|
411
|
+
// them differently from plain strings.
|
|
412
|
+
//
|
|
413
|
+
// IMPORTANT: These are whole-line tokens (pattern + optional
|
|
414
|
+
// horizontal whitespace + newline). Using a whole-line token (like
|
|
415
|
+
// `comment` and `_top_scalar_text`) avoids the ambiguity where the
|
|
416
|
+
// partial integer token `1` wins over the scalar token `1:2:3` via
|
|
417
|
+
// prec — with a whole-line token the match for `1:2:3\n` is length
|
|
418
|
+
// 5 for scalar and no match for integer (`:` breaks the pattern),
|
|
419
|
+
// so scalar correctly wins on non-numeric lines.
|
|
420
|
+
//
|
|
421
|
+
// Float must be tested before integer because the float pattern
|
|
422
|
+
// (decimal point form) is a strict superset of the integer pattern
|
|
423
|
+
// prefix. Float is given prec(3) so it beats integer on `1.5\n`.
|
|
424
|
+
//
|
|
425
|
+
// Integer forms (§ 3.6): hex, octal, binary, decimal with underscores.
|
|
426
|
+
integer: $ => token(prec(2,
|
|
427
|
+
/[+-]?(0x[0-9a-fA-F]([_]?[0-9a-fA-F])*|0o[0-7]([_]?[0-7])*|0b[01]([_]?[01])*|[0-9]([_]?[0-9])*)[ \t]*\r?\n/
|
|
428
|
+
)),
|
|
429
|
+
|
|
430
|
+
// Float forms (§ 3.6): decimal-point form or exponent-only form.
|
|
431
|
+
float: $ => token(prec(3,
|
|
432
|
+
/([+-]?[0-9]([_]?[0-9])*\.[0-9]([_]?[0-9])*([eE][+-]?[0-9]([_]?[0-9])*)?|[+-]?[0-9]([_]?[0-9])*[eE][+-]?[0-9]([_]?[0-9])*)[ \t]*\r?\n/
|
|
433
|
+
)),
|
|
340
434
|
|
|
341
435
|
// ---- Keywords ----
|
|
342
436
|
keyword: $ => seq(
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "tree-sitter-ktav",
|
|
3
|
-
"version": "0.
|
|
3
|
+
"version": "0.6.0",
|
|
4
4
|
"description": "Tree-sitter grammar for Ktav (כְּתָב) — the Written Configuration Format",
|
|
5
5
|
"keywords": [
|
|
6
6
|
"parser",
|
|
@@ -11,7 +11,7 @@
|
|
|
11
11
|
"configuration"
|
|
12
12
|
],
|
|
13
13
|
"author": "Marat K <phpcraftdream@gmail.com>",
|
|
14
|
-
"license": "MIT",
|
|
14
|
+
"license": "MIT OR Apache-2.0",
|
|
15
15
|
"homepage": "https://github.com/ktav-lang/tree-sitter-ktav",
|
|
16
16
|
"repository": {
|
|
17
17
|
"type": "git",
|
package/queries/highlights.scm
CHANGED
|
@@ -14,8 +14,6 @@
|
|
|
14
14
|
; ---- Pair separators ----
|
|
15
15
|
(sep_string) @punctuation.delimiter
|
|
16
16
|
(sep_raw) @punctuation.special
|
|
17
|
-
(sep_int) @punctuation.special
|
|
18
|
-
(sep_float) @punctuation.special
|
|
19
17
|
|
|
20
18
|
; ---- Compound brackets ----
|
|
21
19
|
; The structural openers / closers each form their own visible token
|
|
@@ -43,9 +41,12 @@
|
|
|
43
41
|
(kw_true) @constant.builtin.boolean
|
|
44
42
|
(kw_false) @constant.builtin.boolean
|
|
45
43
|
|
|
44
|
+
; ---- Number literals (new in spec 0.5.0) ----
|
|
45
|
+
(integer) @number
|
|
46
|
+
(float) @number.float
|
|
47
|
+
|
|
46
48
|
; ---- String values ----
|
|
47
|
-
; Plain scalar after `:` —
|
|
48
|
-
; only after `:i`/`:f`; see below.)
|
|
49
|
+
; Plain scalar after `:` — a string.
|
|
49
50
|
(object_pair
|
|
50
51
|
separator: (sep_string)
|
|
51
52
|
value: (scalar) @string)
|
|
@@ -53,33 +54,22 @@
|
|
|
53
54
|
; Raw string after `::`
|
|
54
55
|
(object_pair
|
|
55
56
|
separator: (sep_raw)
|
|
56
|
-
value: (
|
|
57
|
-
|
|
58
|
-
; Typed integer / float bodies
|
|
59
|
-
(object_pair
|
|
60
|
-
separator: (sep_int)
|
|
61
|
-
value: (scalar) @number)
|
|
62
|
-
|
|
63
|
-
(object_pair
|
|
64
|
-
separator: (sep_float)
|
|
65
|
-
value: (scalar) @number.float)
|
|
57
|
+
value: (raw_scalar) @string.special)
|
|
66
58
|
|
|
67
59
|
; ---- Array items ----
|
|
68
60
|
(array_item
|
|
69
61
|
marker: (sep_raw)
|
|
70
|
-
value: (
|
|
71
|
-
|
|
72
|
-
(array_item
|
|
73
|
-
marker: (sep_int)
|
|
74
|
-
value: (scalar) @number)
|
|
75
|
-
|
|
76
|
-
(array_item
|
|
77
|
-
marker: (sep_float)
|
|
78
|
-
value: (scalar) @number.float)
|
|
62
|
+
value: (raw_scalar) @string.special)
|
|
79
63
|
|
|
80
64
|
(array_item
|
|
81
65
|
value: (scalar) @string)
|
|
82
66
|
|
|
67
|
+
; ---- Inline compounds (new in spec 0.5.0) ----
|
|
68
|
+
(inline_object) @punctuation.bracket
|
|
69
|
+
(inline_array) @punctuation.bracket
|
|
70
|
+
(inline_scalar) @string
|
|
71
|
+
(escape_sequence) @string.escape
|
|
72
|
+
|
|
83
73
|
; ---- Multi-line strings ----
|
|
84
74
|
(multiline_stripped) @string
|
|
85
75
|
(multiline_verbatim) @string
|