tree-sitter-ktav 0.3.0 → 0.6.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -171,7 +171,8 @@ rejects (mostly missing-whitespace-after-marker — § 6.10).
171
171
 
172
172
  ## License
173
173
 
174
- MIT — see [LICENSE](LICENSE).
174
+ Dual-licensed under **MIT OR Apache-2.0** — see
175
+ [LICENSE-MIT](LICENSE-MIT) and [LICENSE-APACHE](LICENSE-APACHE).
175
176
 
176
177
  ## Other Ktav repositories
177
178
 
package/README.ru.md CHANGED
@@ -171,7 +171,7 @@ npx tree-sitter test # запускает корпус
171
171
 
172
172
  ## Лицензия
173
173
 
174
- MIT — см. [LICENSE](LICENSE).
174
+ MIT OR Apache-2.0. См. [LICENSE-MIT](LICENSE-MIT) и [LICENSE-APACHE](LICENSE-APACHE).
175
175
 
176
176
  ## Другие репозитории Ktav
177
177
 
package/README.zh.md CHANGED
@@ -165,7 +165,7 @@ CLI 即可构建。
165
165
 
166
166
  ## 许可证
167
167
 
168
- MIT — 见 [LICENSE](LICENSE)。
168
+ MIT OR Apache-2.0。详见 [LICENSE-MIT](LICENSE-MIT) 和 [LICENSE-APACHE](LICENSE-APACHE)。
169
169
 
170
170
  ## 其他 Ktav 仓库
171
171
 
package/grammar.js CHANGED
@@ -1,24 +1,43 @@
1
1
  /**
2
2
  * Tree-sitter grammar for Ktav (כְּתָב) — the Written Configuration Format.
3
3
  *
4
- * Spec: https://github.com/ktav-lang/spec/blob/main/versions/0.1/spec.md
4
+ * Spec: https://github.com/ktav-lang/spec/blob/main/versions/0.6/spec.md
5
5
  *
6
6
  * Ktav is line-oriented. Every line is one of:
7
7
  * - blank
8
- * - a comment (`# ...`)
9
- * - a key:value pair (with markers `:`, `::`, `:i`, `:f`)
8
+ * - a comment (`## ...`) ← NOTE: double-hash in 0.5.0; single `#` is content
9
+ * - a key:value pair (with markers `:` or `::`)
10
10
  * - a structural opener / closer for compounds (`{`, `}`, `[`, `]`,
11
11
  * `(`, `((`, `)`, `))`)
12
- * - an array item (inside an open `[` array)
12
+ * - an array item (inside an open `[` array, or at the top level)
13
13
  * - raw content of a multi-line string
14
14
  *
15
+ * Changes from 0.5.0 to 0.6.0:
16
+ * - Keys now process escape sequences (§ 3.7). The escape table grows
17
+ * to 10 entries by adding `\.` (literal dot — does NOT split the
18
+ * dotted path) and `\:` (literal colon — does NOT act as the
19
+ * key/value separator). `\` becomes the escape lead in keys.
20
+ * Breaking: a raw `\` in a key now requires `\\`.
21
+ *
22
+ * Changes from 0.3.0 (spec 0.1.1) to 0.5.0:
23
+ * - Comment marker changed from `#` to `##`. Single `#` is now a
24
+ * content byte (allowed in keys and scalar values).
25
+ * - Typed markers `:i` and `:f` removed. Only `:` and `::` remain.
26
+ * - Inline compounds: `{key: value, ...}` and `[v1, v2, ...]` are
27
+ * now valid as pair values or array items (inline_object /
28
+ * inline_array rules).
29
+ * - Number literals: hex (`0x`), octal (`0o`), binary (`0b`), decimal
30
+ * with underscore separators, and floats with `.` or exponent.
31
+ * These are captured as distinct node kinds for highlighting.
32
+ * - Escape sequences inside inline scalars: `\\`, `\,`, `\}`, `\]`,
33
+ * `\{`, `\[`, `\n`, `\r`.
34
+ *
15
35
  * Strategy:
16
36
  * - Newlines are explicit (`_newline`) and structural openers/closers
17
37
  * are tokens that include the trailing whitespace + newline so they
18
38
  * can NEVER be confused with a scalar starting with the same byte.
19
- * - The four pair separators (`:`, `::`, `:i`, `:f`) are recognized
20
- * by the lexer with longest-match precedence (`::` > `:`, `:i`/`:f`
21
- * > `:`).
39
+ * - The two pair separators (`:`, `::`) are recognized by the lexer
40
+ * with longest-match precedence (`::` > `:`).
22
41
  * - The mandatory-whitespace-after-marker rule (§ 6.10) is enforced
23
42
  * by an external scanner token `_marker_ws`, which is a zero-width
24
43
  * assertion that only succeeds when the byte right after the
@@ -31,6 +50,8 @@
31
50
  * a parse error.
32
51
  * - Multi-line string content is captured as a sequence of opaque
33
52
  * "raw lines" up to the matching terminator.
53
+ * - Inline compound values are fully parsed; escape sequences inside
54
+ * inline scalars are captured as `escape_sequence` nodes.
34
55
  * - Indentation is not significant (matches the spec).
35
56
  */
36
57
 
@@ -56,29 +77,12 @@ module.exports = grammar({
56
77
 
57
78
  rules: {
58
79
  // The top-level document is a sequence of lines. Per spec § 5.0.1
59
- // (added in spec 0.1.1) the root may be either an Object (the
60
- // historical default — every content line is a pair) or an Array
61
- // (every content line is an array item). Tree-sitter is a generic
62
- // parser of trees: rather than encode the spec's "first content
63
- // line decides" rule structurally, we accept either kind of line
64
- // anywhere at the top level. The reference parser performs the
65
- // semantic § 5.0.1 dispatch and rejects mixed roots; the grammar's
66
- // job here is only to produce a clean syntax tree for both shapes
67
- // (and for editors / LSPs) without spurious ERROR nodes on bare
68
- // top-level scalars or `:i`/`:f`/`::` lines.
80
+ // the root may be either an Object or an Array. Tree-sitter accepts
81
+ // both kinds of line anywhere; semantic dispatch is left to the
82
+ // reference parser.
69
83
  source_file: $ => repeat($._line),
70
84
 
71
85
  // ---- Top-level lines ----
72
- //
73
- // `top_array_item` covers the top-level Array case (§ 5.0.1): a
74
- // bare scalar, a typed-marker item (`:: …`, `:i …`, `:f …`), a
75
- // lone `{` / `[` opener, or a `(` / `((` multi-line opener. It
76
- // is structurally identical to `array_item` (used inside `[ ]`)
77
- // but is kept as a distinct rule so we can apply a lexer-level
78
- // disambiguation: a line that classifies as a pair (§ 5.3) must
79
- // remain a pair at the top level (spec § 5.0.1 step 2), so
80
- // `top_array_item`'s plain bare-scalar branch is given lower
81
- // precedence than `object_pair` via `prec.dynamic`.
82
86
  _line: $ => choice(
83
87
  $.comment,
84
88
  $.blank_line,
@@ -92,27 +96,25 @@ module.exports = grammar({
92
96
 
93
97
  // ---- Comment ----
94
98
  //
95
- // Captured as a single whole-line token so it is the
96
- // unambiguous longest match at line start when the line
97
- // begins with `#`. Without this, `_top_scalar_text` (also a
98
- // whole-line token) would tie or beat a structurally-defined
99
- // comment that's split into `'#'` + tail + newline pieces.
100
- comment: $ => token(prec(1, /#[^\r\n]*\r?\n/)),
99
+ // In spec 0.5.0 a comment starts with `##` (two hashes). A single
100
+ // `#` is ordinary content. The token captures the whole line
101
+ // including the trailing newline to beat `_top_scalar_text` at the
102
+ // lexer's longest-match step.
103
+ comment: $ => token(prec(1, /##[^\r\n]*\r?\n/)),
101
104
 
102
105
  // ---- Object pair ----
103
106
  //
104
107
  // After the separator, the external `_marker_ws` token asserts
105
- // that the next byte is whitespace, CR, LF, or EOF. This enforces
106
- // § 6.10 (mandatory whitespace after marker): `key:value` fails.
108
+ // that the next byte is whitespace, CR, LF, or EOF (§ 6.10).
107
109
  object_pair: $ => choice(
108
- // After `::`, `:i`, `:f` the body is a literal/typed scalar — § 5.2:
109
- // it is NOT dispatched through compound-opener / multi-line dispatch.
110
- // Only `empty_value` (immediate newline) or `scalar` is allowed.
110
+ // After `::` the body is a literal scalar (NOT dispatched through
111
+ // compound-opener / multi-line dispatch). `raw_scalar` accepts any
112
+ // line content including `(`, `{`, `[`-starting text.
111
113
  seq(
112
114
  field('key', $.key),
113
- field('separator', choice($.sep_raw, $.sep_int, $.sep_float)),
115
+ field('separator', $.sep_raw),
114
116
  $._marker_ws,
115
- field('value', choice($.empty_value, $.scalar)),
117
+ field('value', choice($.empty_value, $.raw_scalar)),
116
118
  ),
117
119
  // After `:` the body goes through the full § 5.2 dispatch.
118
120
  seq(
@@ -126,27 +128,39 @@ module.exports = grammar({
126
128
  ),
127
129
  ),
128
130
 
129
- // Tree-sitter prefers longest token match. To make sure `::` wins
130
- // over `:`, `::` is given higher precedence; same for `:i`/`:f`.
131
+ // `::` wins over `:` via higher precedence.
131
132
  sep_raw: $ => token(prec(3, '::')),
132
- sep_int: $ => token(prec(2, ':i')),
133
- sep_float: $ => token(prec(2, ':f')),
134
133
  sep_string: $ => token(prec(1, ':')),
135
134
 
136
135
  // ---- Keys ----
137
136
  key: $ => choice(
138
- $._key_segment,
137
+ $._spaced_key,
139
138
  $.dotted_key,
140
139
  ),
141
140
 
141
+ // A key may contain internal whitespace (spec 0.5.0 § 4): the run of
142
+ // space-separated segments up to the separator is one key
143
+ // (`multi word key: value`). The inter-segment spaces are `extras`,
144
+ // so the `key` node still spans the whole text with no named children
145
+ // (renders as `(key)`, same as a single-segment key).
146
+ _spaced_key: $ => prec.left(repeat1($._key_segment)),
147
+
142
148
  dotted_key: $ => prec.left(seq(
143
149
  $._key_segment,
144
150
  repeat1(seq('.', $._key_segment)),
145
151
  )),
146
152
 
147
- // Key segment: any chars except whitespace, "[", "]", "{", "}",
148
- // ":", "#", ".".
149
- _key_segment: $ => /[^\s\[\]\{\}:#.\r\n]+/,
153
+ // Key segment (spec 0.6.0 § 4): a non-empty run of plain key bytes
154
+ // and/or escape sequences. Plain key bytes exclude whitespace, the
155
+ // bracket/paren/brace bytes, `:`, `,`, the dotted-path separator `.`,
156
+ // `#` (reserved for comment marker `##`), and the escape lead `\`.
157
+ // Those structural bytes — including `\.` and `\:` — can now appear
158
+ // inside a key when escaped (§ 3.7, expanded in 0.6.0 to include
159
+ // `\.` and `\:`).
160
+ //
161
+ // The dotted-path separator is an UNescaped `.`; an unescaped `\`
162
+ // is always the start of an escape sequence (10 forms in 0.6.0).
163
+ _key_segment: $ => /([^\s\[\]\{\}\(\):#,.\r\n\\]|\\[\\,\}\]\{\[nr.:])+/,
150
164
 
151
165
  // ---- Value line ----
152
166
  _value_line: $ => choice(
@@ -160,8 +174,14 @@ module.exports = grammar({
160
174
  $.empty_array,
161
175
  $.empty_paren,
162
176
  $.empty_double_paren,
177
+ // Inline compounds (new in 0.5.0).
178
+ $.inline_object,
179
+ $.inline_array,
163
180
  // Keywords (single token followed by newline).
164
181
  $.keyword,
182
+ // Number literals (distinct nodes for highlighting).
183
+ $.integer,
184
+ $.float,
165
185
  // Scalar — catch-all line content.
166
186
  $.scalar,
167
187
  ),
@@ -176,25 +196,12 @@ module.exports = grammar({
176
196
  empty_double_paren: $ => seq(token(prec(5, '(())')), $._newline),
177
197
 
178
198
  // ---- Multi-line compounds ----
179
- //
180
- // Openers are tokens that include the rest of the line up to and
181
- // including the newline, so they cannot be confused with a scalar
182
- // starting with the same byte. Closers, by contrast, are split
183
- // into the bracket character(s) plus the external `_strict_eol`
184
- // token; this lets the scanner reject pathological lines like
185
- // `} trailing_text\n` (§ 5.6.1 and the cleanliness rule for
186
- // object/array closers).
187
199
  open_brace: $ => token(prec(4, /\{[ \t]*\r?\n/)),
188
200
  close_brace: $ => seq(token(prec(4, '}')), $._strict_eol),
189
201
  open_bracket: $ => token(prec(4, /\[[ \t]*\r?\n/)),
190
202
  close_bracket: $ => seq(token(prec(4, ']')), $._strict_eol),
191
203
  open_paren: $ => token(prec(4, /\([ \t]*\r?\n/)),
192
204
  open_dparen: $ => token(prec(5, /\(\([ \t]*\r?\n/)),
193
- // close_paren / close_dparen are emitted by the external scanner as
194
- // `_stripped_close` / `_verbatim_close`. They are context-sensitive:
195
- // inside `(...)` only `)` closes (a `))` line is content); inside
196
- // `((...))` only `))` closes (a single `)` line is content). The
197
- // scanner consults `valid_symbols` to decide which form to attempt.
198
205
  close_paren: $ => $._stripped_close,
199
206
  close_dparen: $ => $._verbatim_close,
200
207
 
@@ -210,62 +217,128 @@ module.exports = grammar({
210
217
  $.close_bracket,
211
218
  ),
212
219
 
213
- // ---- Array items ----
220
+ // ---- Inline compounds (new in spec 0.5.0) ----
214
221
  //
215
- // Marker-prefixed items must, like pair lines, have whitespace or
216
- // EOL after the marker (§ 6.10).
222
+ // `{key: value, key2: value2}` and `[v1, v2, v3]` are valid as a
223
+ // value on the right-hand side of a pair or as an array item.
224
+ // Trailing comma is allowed. Nesting is supported.
225
+ //
226
+ // These are followed by a newline (they consume the rest of the line).
227
+ inline_object: $ => seq(
228
+ '{',
229
+ optional($._inline_pair_list),
230
+ '}',
231
+ $._newline,
232
+ ),
233
+
234
+ inline_array: $ => seq(
235
+ '[',
236
+ optional($._inline_item_list),
237
+ ']',
238
+ $._newline,
239
+ ),
240
+
241
+ _inline_pair_list: $ => seq(
242
+ $.inline_pair,
243
+ repeat(seq(',', $.inline_pair)),
244
+ optional(','),
245
+ ),
246
+
247
+ // The value is optional: `{x:, y: 1}` and `{empty:}` are valid —
248
+ // a separator immediately followed by `,` or `}` is an empty value
249
+ // (spec 0.5.0 § 5.8).
250
+ inline_pair: $ => seq(
251
+ field('key', $.key),
252
+ field('separator', choice($.sep_raw, $.sep_string)),
253
+ optional(field('value', $.inline_value)),
254
+ ),
255
+
256
+ _inline_item_list: $ => seq(
257
+ $.inline_value,
258
+ repeat(seq(',', $.inline_value)),
259
+ optional(','),
260
+ ),
261
+
262
+ // An inline value is either a nested inline compound, or an inline
263
+ // scalar (which may contain escape sequences).
264
+ inline_value: $ => choice(
265
+ $.nested_inline_object,
266
+ $.nested_inline_array,
267
+ $.inline_scalar,
268
+ ),
269
+
270
+ nested_inline_object: $ => seq(
271
+ '{',
272
+ optional($._inline_pair_list),
273
+ '}',
274
+ ),
275
+
276
+ nested_inline_array: $ => seq(
277
+ '[',
278
+ optional($._inline_item_list),
279
+ ']',
280
+ ),
281
+
282
+ // An inline scalar is terminated by an unescaped `,`, `}`, or `]`.
283
+ // It may contain escape sequences (§ 3.7).
284
+ //
285
+ // The first chunk must NOT begin with `{` or `[`: a value position
286
+ // that opens with `{`/`[` is a nested compound, not a scalar. After
287
+ // the first character those bytes are ordinary literal content
288
+ // (`hello{world`, `mid[bracket`). This head/rest split keeps
289
+ // `nested_inline_*` vs `inline_scalar` unambiguous without sacrificing
290
+ // mid-value literal braces (§ 5.8).
291
+ inline_scalar: $ => seq(
292
+ choice($.escape_sequence, $._inline_scalar_head),
293
+ repeat(choice($.escape_sequence, $._inline_scalar_text)),
294
+ ),
295
+
296
+ // Escape sequences recognised inside inline scalars (§ 3.7).
297
+ // In 0.6.0 the table grows to 10 forms by adding `\.` and `\:` —
298
+ // they yield literal `.` and `:` respectively. Inside a value these
299
+ // are redundant (a value's `.`/`:` is already a literal byte) but
300
+ // remain accepted for symmetry with key parsing.
301
+ escape_sequence: $ => token(choice(
302
+ '\\\\',
303
+ '\\,',
304
+ '\\}',
305
+ '\\]',
306
+ '\\{',
307
+ '\\[',
308
+ '\\n',
309
+ '\\r',
310
+ '\\.',
311
+ '\\:',
312
+ )),
313
+
314
+ // Leading chunk of a scalar: first byte excludes whitespace and the
315
+ // openers `{`/`[` (so a value starting with an opener is a nested
316
+ // compound), plus the usual `\` `,` `}` `]` / CR / LF. Subsequent
317
+ // bytes allow `{`/`[` as literal content.
318
+ _inline_scalar_head: $ => token(/[^\s\\,\{\[\}\]\r\n][^\\,\}\]\r\n]*/),
319
+
320
+ // Continuation text after the head (or after an escape): any byte
321
+ // except `\`, `,`, the closers `}` `]`, and CR / LF. Open delimiters
322
+ // `{`/`[` are allowed here as literal content.
323
+ _inline_scalar_text: $ => token(/[^\\,\}\]\r\n]+/),
324
+
325
+ // ---- Array items ----
217
326
  array_item: $ => choice(
218
327
  seq(
219
328
  field('marker', $.sep_raw),
220
329
  $._marker_ws,
221
- field('value', choice($.empty_value, $.scalar)),
222
- ),
223
- seq(
224
- field('marker', $.sep_int),
225
- $._marker_ws,
226
- field('value', choice($.empty_value, $.scalar)),
227
- ),
228
- seq(
229
- field('marker', $.sep_float),
230
- $._marker_ws,
231
- field('value', choice($.empty_value, $.scalar)),
330
+ field('value', choice($.empty_value, $.raw_scalar)),
232
331
  ),
233
332
  // Plain value item — same set as object pair value.
234
333
  field('value', $._value_line),
235
334
  ),
236
335
 
237
- // Top-level array item (§ 5.0.1, added in spec 0.1.1).
238
- //
239
- // Structurally a sibling of `array_item` (used inside `[ ]`),
240
- // but with two differences that realise spec § 5.0.1 step 2:
241
- //
242
- // * The plain bare-scalar branch uses `_top_scalar_text`, a
243
- // token that disallows `:` anywhere in the line. This makes
244
- // any line containing a `:` (i.e. any pair-shaped line)
245
- // parseable ONLY as `object_pair`, never as a top-level
246
- // bare-scalar — which is what § 5.0.1 step 2 requires.
247
- // Per spec note: to force a colon-bearing scalar at the root
248
- // to be an Array item, use the raw marker (`:: foo: bar`).
249
- //
250
- // * Compound openers (`{`, `[`, `(`, `((`) and the empty-inline
251
- // compound forms (`{}` / `[]` / `()` / `(())`) and the
252
- // keywords are still allowed unchanged — they are
253
- // unambiguous lines that cannot be confused with a pair.
336
+ // Top-level array item (§ 5.0.1).
254
337
  top_array_item: $ => choice(
255
338
  seq(
256
339
  field('marker', $.sep_raw),
257
340
  $._marker_ws,
258
- field('value', choice($.empty_value, $.scalar)),
259
- ),
260
- seq(
261
- field('marker', $.sep_int),
262
- $._marker_ws,
263
- field('value', choice($.empty_value, $.scalar)),
264
- ),
265
- seq(
266
- field('marker', $.sep_float),
267
- $._marker_ws,
268
- field('value', choice($.empty_value, $.scalar)),
341
+ field('value', choice($.empty_value, $.raw_scalar)),
269
342
  ),
270
343
  field('value', $.compound_object),
271
344
  field('value', $.compound_array),
@@ -275,35 +348,18 @@ module.exports = grammar({
275
348
  field('value', $.empty_array),
276
349
  field('value', $.empty_paren),
277
350
  field('value', $.empty_double_paren),
351
+ field('value', $.inline_object),
352
+ field('value', $.inline_array),
278
353
  field('value', $.keyword),
279
- // Bare-scalar branch — the line is captured as a `scalar`
280
- // node whose text comes from the `_top_scalar_text` token
281
- // (no `:` allowed). Wrapping in `top_scalar` keeps the AST
282
- // tidy: tools see exactly the same `(scalar)` shape as
283
- // elsewhere, with the only difference being which lexer
284
- // rule produced the inner token.
354
+ field('value', $.integer),
355
+ field('value', $.float),
285
356
  field('value', $.top_scalar),
286
357
  ),
287
358
 
288
- // `top_scalar` is a one-token whole-line scalar used at the
289
- // document root for spec § 5.0.1's bare-scalar Array-item
290
- // case. The token swallows the trailing newline, which makes
291
- // it strictly longer than the `_key_segment` token a pair
292
- // would emit (`_key_segment` stops at the first `:` / `\n` /
293
- // structural byte). Two consequences:
294
- //
295
- // 1. On a colon-free line such as `foo\n`, the whole-line
296
- // `_top_scalar_text` token (length 4) wins the lexer's
297
- // longest-match against `_key_segment` (length 3) — so
298
- // the parser commits to the top-level Array-item path,
299
- // not the always-ERROR pair-without-separator path.
300
- // 2. On a pair-shaped line such as `name: Russia\n`, the
301
- // `_top_scalar_text` regex CANNOT match (it forbids
302
- // `:`), so `_key_segment` is the only viable token at
303
- // line-start and the parser correctly enters
304
- // `object_pair` (spec § 5.0.1 step 2).
359
+ // `top_scalar` — bare-scalar at the document root. Forbids `:` so
360
+ // that pair-shaped lines always parse as `object_pair`.
305
361
  top_scalar: $ => $._top_scalar_text,
306
- _top_scalar_text: $ => token(/[^\s:\r\n][^:\r\n]*\r?\n/),
362
+ _top_scalar_text: $ => token(/[^\s:\{\[\(\r\n][^:\r\n]*\r?\n/),
307
363
 
308
364
  // ---- Multi-line strings ----
309
365
  multiline_stripped: $ => seq(
@@ -318,11 +374,6 @@ module.exports = grammar({
318
374
  $.close_dparen,
319
375
  ),
320
376
 
321
- // A content line inside a multi-line string. Captured as one
322
- // token. The closer tokens (`)`, `))` plus _strict_eol) win at
323
- // the LR(1) boundary because they require their own line; any
324
- // line whose content includes more than just the terminator
325
- // falls through to this rule.
326
377
  multiline_content_line: $ => token(prec(-1, /[^\r\n]*\r?\n/)),
327
378
 
328
379
  // ---- Scalar (default value body, until end of line) ----
@@ -331,12 +382,55 @@ module.exports = grammar({
331
382
  $._newline,
332
383
  ),
333
384
 
385
+ // `raw_scalar` is used exclusively after `::` (raw marker). It accepts
386
+ // any non-empty line content, including `(`, `{`, `[`-starting text.
387
+ // Per spec § 5.2: after `::` the body is NEVER dispatched as a
388
+ // compound opener or multi-line string opener — it is always a literal
389
+ // String value. This is a separate rule (not `scalar`) because
390
+ // `scalar`'s underlying `_scalar_text` deliberately excludes those
391
+ // opening bytes to avoid lexer ambiguity in the non-raw value context.
392
+ raw_scalar: $ => seq(
393
+ $._raw_scalar_text,
394
+ $._newline,
395
+ ),
396
+ _raw_scalar_text: $ => /[^\s\r\n][^\r\n]*/,
397
+
334
398
  // Scalar text: any non-whitespace, non-newline content up to end
335
- // of line. `#` is allowed (`color: #ff00ff` is a valid value);
336
- // the `#` only opens a comment when it is the first non-whitespace
337
- // char of a line — tree-sitter's context-aware lexer disambiguates
338
- // because `comment` is only valid at a line-start parse state.
339
- _scalar_text: $ => /[^\s\r\n][^\r\n]*/,
399
+ // of line. Both `#` and `##` are allowed as content bytes in 0.5.0.
400
+ // Lines starting with `{` or `[` are always handled by structural
401
+ // rules (compound_object, compound_array, empty_object, empty_array,
402
+ // inline_object, inline_array), so `_scalar_text` explicitly excludes
403
+ // those opening bytes at position 0 to avoid the greedy-token
404
+ // ambiguity. Lines starting with `(` are handled by multiline or
405
+ // empty-paren rules likewise.
406
+ _scalar_text: $ => /[^\s\{\[\(\r\n][^\r\n]*/,
407
+
408
+ // ---- Number literals ----
409
+ //
410
+ // Captured as distinct node kinds so syntax highlighters can colour
411
+ // them differently from plain strings.
412
+ //
413
+ // IMPORTANT: These are whole-line tokens (pattern + optional
414
+ // horizontal whitespace + newline). Using a whole-line token (like
415
+ // `comment` and `_top_scalar_text`) avoids the ambiguity where the
416
+ // partial integer token `1` wins over the scalar token `1:2:3` via
417
+ // prec — with a whole-line token the match for `1:2:3\n` is length
418
+ // 5 for scalar and no match for integer (`:` breaks the pattern),
419
+ // so scalar correctly wins on non-numeric lines.
420
+ //
421
+ // Float must be tested before integer because the float pattern
422
+ // (decimal point form) is a strict superset of the integer pattern
423
+ // prefix. Float is given prec(3) so it beats integer on `1.5\n`.
424
+ //
425
+ // Integer forms (§ 3.6): hex, octal, binary, decimal with underscores.
426
+ integer: $ => token(prec(2,
427
+ /[+-]?(0x[0-9a-fA-F]([_]?[0-9a-fA-F])*|0o[0-7]([_]?[0-7])*|0b[01]([_]?[01])*|[0-9]([_]?[0-9])*)[ \t]*\r?\n/
428
+ )),
429
+
430
+ // Float forms (§ 3.6): decimal-point form or exponent-only form.
431
+ float: $ => token(prec(3,
432
+ /([+-]?[0-9]([_]?[0-9])*\.[0-9]([_]?[0-9])*([eE][+-]?[0-9]([_]?[0-9])*)?|[+-]?[0-9]([_]?[0-9])*[eE][+-]?[0-9]([_]?[0-9])*)[ \t]*\r?\n/
433
+ )),
340
434
 
341
435
  // ---- Keywords ----
342
436
  keyword: $ => seq(
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "tree-sitter-ktav",
3
- "version": "0.3.0",
3
+ "version": "0.6.0",
4
4
  "description": "Tree-sitter grammar for Ktav (כְּתָב) — the Written Configuration Format",
5
5
  "keywords": [
6
6
  "parser",
@@ -11,7 +11,7 @@
11
11
  "configuration"
12
12
  ],
13
13
  "author": "Marat K <phpcraftdream@gmail.com>",
14
- "license": "MIT",
14
+ "license": "MIT OR Apache-2.0",
15
15
  "homepage": "https://github.com/ktav-lang/tree-sitter-ktav",
16
16
  "repository": {
17
17
  "type": "git",
@@ -14,8 +14,6 @@
14
14
  ; ---- Pair separators ----
15
15
  (sep_string) @punctuation.delimiter
16
16
  (sep_raw) @punctuation.special
17
- (sep_int) @punctuation.special
18
- (sep_float) @punctuation.special
19
17
 
20
18
  ; ---- Compound brackets ----
21
19
  ; The structural openers / closers each form their own visible token
@@ -43,9 +41,12 @@
43
41
  (kw_true) @constant.builtin.boolean
44
42
  (kw_false) @constant.builtin.boolean
45
43
 
44
+ ; ---- Number literals (new in spec 0.5.0) ----
45
+ (integer) @number
46
+ (float) @number.float
47
+
46
48
  ; ---- String values ----
47
- ; Plain scalar after `:` — usually a string. (Numeric coercion happens
48
- ; only after `:i`/`:f`; see below.)
49
+ ; Plain scalar after `:` — a string.
49
50
  (object_pair
50
51
  separator: (sep_string)
51
52
  value: (scalar) @string)
@@ -53,33 +54,22 @@
53
54
  ; Raw string after `::`
54
55
  (object_pair
55
56
  separator: (sep_raw)
56
- value: (scalar) @string.special)
57
-
58
- ; Typed integer / float bodies
59
- (object_pair
60
- separator: (sep_int)
61
- value: (scalar) @number)
62
-
63
- (object_pair
64
- separator: (sep_float)
65
- value: (scalar) @number.float)
57
+ value: (raw_scalar) @string.special)
66
58
 
67
59
  ; ---- Array items ----
68
60
  (array_item
69
61
  marker: (sep_raw)
70
- value: (scalar) @string.special)
71
-
72
- (array_item
73
- marker: (sep_int)
74
- value: (scalar) @number)
75
-
76
- (array_item
77
- marker: (sep_float)
78
- value: (scalar) @number.float)
62
+ value: (raw_scalar) @string.special)
79
63
 
80
64
  (array_item
81
65
  value: (scalar) @string)
82
66
 
67
+ ; ---- Inline compounds (new in spec 0.5.0) ----
68
+ (inline_object) @punctuation.bracket
69
+ (inline_array) @punctuation.bracket
70
+ (inline_scalar) @string
71
+ (escape_sequence) @string.escape
72
+
83
73
  ; ---- Multi-line strings ----
84
74
  (multiline_stripped) @string
85
75
  (multiline_verbatim) @string