lilt 0.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +7 -0
- data/LICENSE.txt +21 -0
- data/README.md +265 -0
- data/lib/lilt/lexer.rb +42 -0
- data/lib/lilt/parser.rb +104 -0
- data/lib/lilt/sexp.rb +54 -0
- data/lib/lilt/version.rb +5 -0
- data/lib/lilt.rb +6 -0
- metadata +52 -0
checksums.yaml
ADDED
|
@@ -0,0 +1,7 @@
|
|
|
1
|
+
---
|
|
2
|
+
SHA256:
|
|
3
|
+
metadata.gz: 27310c339a7875bd8f1e6705459689c05dc978c0c5dc6e1d59534b650493528e
|
|
4
|
+
data.tar.gz: 34cb476a1f8e961d55af835aeb5e490b2fbb7dde94d3b58fe3647c90181b40b2
|
|
5
|
+
SHA512:
|
|
6
|
+
metadata.gz: 7159ca1d750e56ee6c0bbb8c958db5a22188dd8784895a8c9765a5f9038ec8612258914ca0cd6786e003d1f7883f99c3d0de020aa72fc7594f9fdf3324ca0ed4
|
|
7
|
+
data.tar.gz: 4fbb86254219e96138bf7d8409067791dcf3e4d5a43e0e1bc3b595969cfc263b6f2e60ff7a9aff9947cbee5e508dba8311ea44bc632eba5137fc80556daac37a
|
data/LICENSE.txt
ADDED
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
The MIT License (MIT)
|
|
2
|
+
|
|
3
|
+
Copyright (c) 2026 Phil Crissman
|
|
4
|
+
|
|
5
|
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
|
6
|
+
of this software and associated documentation files (the "Software"), to deal
|
|
7
|
+
in the Software without restriction, including without limitation the rights
|
|
8
|
+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
|
9
|
+
copies of the Software, and to permit persons to whom the Software is
|
|
10
|
+
furnished to do so, subject to the following conditions:
|
|
11
|
+
|
|
12
|
+
The above copyright notice and this permission notice shall be included in
|
|
13
|
+
all copies or substantial portions of the Software.
|
|
14
|
+
|
|
15
|
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
|
16
|
+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
|
17
|
+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
|
18
|
+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
|
19
|
+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
|
20
|
+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN
|
|
21
|
+
THE SOFTWARE.
|
data/README.md
ADDED
|
@@ -0,0 +1,265 @@
|
|
|
1
|
+
# Lilt
|
|
2
|
+
|
|
3
|
+
A **Li**ttle **L**anguage **T**oolkit.
|
|
4
|
+
|
|
5
|
+
I've been writing a lot of small languages, mostly lambda calculi and similar experiments,
|
|
6
|
+
from [TAPL](https://www.cis.upenn.edu/~bcpierce/tapl/) or other sources. I started to notice that the lexer and parser were all generally
|
|
7
|
+
the same, and were not the most interesting part of the process, so I thought I would try
|
|
8
|
+
to create a library to make it easier to get a lexer and parser going for small languages.
|
|
9
|
+
|
|
10
|
+
It has:
|
|
11
|
+
|
|
12
|
+
- **A table-driven lexer.** You write the rules for recognizing tokens, and get back the
|
|
13
|
+
list of tokens with line and column numbers.
|
|
14
|
+
- **A Pratt parser engine.** You write tables of handlers, and get back whatever AST
|
|
15
|
+
your handlers build.
|
|
16
|
+
- **An s-expression reader and printer.** It gives you a second front end for free, and
|
|
17
|
+
a readable way to print any AST.
|
|
18
|
+
|
|
19
|
+
Grammars are plain data (arrays and hashes of procs), and the AST is yours: Lilt
|
|
20
|
+
doesn't define node types. I've been using `Data.define` for nodes, which gives you
|
|
21
|
+
structural equality and pattern matching for free.
|
|
22
|
+
|
|
23
|
+
## Installation
|
|
24
|
+
|
|
25
|
+
Add it to your Gemfile:
|
|
26
|
+
|
|
27
|
+
```ruby
|
|
28
|
+
gem "lilt"
|
|
29
|
+
```
|
|
30
|
+
|
|
31
|
+
or install it directly:
|
|
32
|
+
|
|
33
|
+
```sh
|
|
34
|
+
gem install lilt
|
|
35
|
+
```
|
|
36
|
+
|
|
37
|
+
It requires Ruby 3.2 or newer (for `Data`).
|
|
38
|
+
|
|
39
|
+
## Lexing
|
|
40
|
+
|
|
41
|
+
A lexer is a list of rules, each a pattern and a token type. Rules are tried in order,
|
|
42
|
+
and the first match wins. A string matches literally; a regex matches a pattern; a
|
|
43
|
+
`nil` type means "skip this".
|
|
44
|
+
|
|
45
|
+
```ruby
|
|
46
|
+
require "lilt"
|
|
47
|
+
|
|
48
|
+
RULES = [
|
|
49
|
+
[/\s+/, nil], # skip whitespace
|
|
50
|
+
["+", :plus],
|
|
51
|
+
["*", :star],
|
|
52
|
+
["(", :lparen],
|
|
53
|
+
[")", :rparen],
|
|
54
|
+
[/\d+/, :int],
|
|
55
|
+
[/[a-z]\w*/, :ident],
|
|
56
|
+
]
|
|
57
|
+
|
|
58
|
+
Lilt::Lexer.lex("let x", RULES, keywords: %w[let])
|
|
59
|
+
# => [#<data Lilt::Token type=:let, value="let", line=1, col=1>,
|
|
60
|
+
# #<data Lilt::Token type=:ident, value="x", line=1, col=5>,
|
|
61
|
+
# #<data Lilt::Token type=:eof, value=nil, line=1, col=6>]
|
|
62
|
+
```
|
|
63
|
+
|
|
64
|
+
`keywords:` promotes any token whose whole text is a keyword to its own type, so `let`
|
|
65
|
+
becomes `:let` while `letter` stays an `:ident`. Every token list ends with `:eof`.
|
|
66
|
+
Columns count characters, not bytes, so `λ` is fine. Input that no rule matches raises
|
|
67
|
+
`Lilt::LexError` with its position.
|
|
68
|
+
|
|
69
|
+
## Parsing
|
|
70
|
+
|
|
71
|
+
A grammar table has `prefix:` handlers (for tokens that start an expression) and
|
|
72
|
+
`infix:` entries (for tokens that continue one). An infix entry is a binding power and a
|
|
73
|
+
handler; `Lilt.binary` builds the common case for you.
|
|
74
|
+
|
|
75
|
+
```ruby
|
|
76
|
+
Num = Data.define(:value)
|
|
77
|
+
Add = Data.define(:left, :right)
|
|
78
|
+
Mul = Data.define(:left, :right)
|
|
79
|
+
|
|
80
|
+
ARITH = {
|
|
81
|
+
prefix: {
|
|
82
|
+
int: proc { |tok| Num.new(tok.value.to_i) },
|
|
83
|
+
lparen: Lilt.group(:rparen),
|
|
84
|
+
},
|
|
85
|
+
infix: {
|
|
86
|
+
plus: Lilt.binary(10) { |l, r| Add.new(l, r) },
|
|
87
|
+
star: Lilt.binary(20) { |l, r| Mul.new(l, r) },
|
|
88
|
+
},
|
|
89
|
+
}
|
|
90
|
+
|
|
91
|
+
ast = Lilt.parse(ARITH, Lilt::Lexer.lex("1 + 2 * 3", RULES))
|
|
92
|
+
# => #<data Add left=#<data Num value=1>, right=#<data Mul left=#<data Num value=2>, right=#<data Num value=3>>>
|
|
93
|
+
|
|
94
|
+
Lilt::Sexp.print(ast)
|
|
95
|
+
# => "(add (num 1) (mul (num 2) (num 3)))"
|
|
96
|
+
```
|
|
97
|
+
|
|
98
|
+
A higher binding power binds more tightly. `Lilt.binary(bp)` is left-associative;
|
|
99
|
+
`Lilt.binary(bp, :right)` is right-associative. `Lilt.group(:rparen)` parses one
|
|
100
|
+
expression and then expects the closing token.
|
|
101
|
+
|
|
102
|
+
For prefix operators like unary minus, `Lilt.prefix(bp) { |operand| Neg.new(operand) }`
|
|
103
|
+
parses its operand at `bp`. With a high `bp` (say 70), `-a * b` means `(-a) * b`; with
|
|
104
|
+
one between `*` and `^`, `-a ^ b` means `-(a ^ b)`. A token can have both a prefix and
|
|
105
|
+
an infix handler, so `-` can mean negation and subtraction.
|
|
106
|
+
|
|
107
|
+
Since the AST is just data, doing something with it is a `case`/`in`:
|
|
108
|
+
|
|
109
|
+
```ruby
|
|
110
|
+
def calc(e)
|
|
111
|
+
case e
|
|
112
|
+
in Num[n] then n
|
|
113
|
+
in Add[l, r] then calc(l) + calc(r)
|
|
114
|
+
in Mul[l, r] then calc(l) * calc(r)
|
|
115
|
+
end
|
|
116
|
+
end
|
|
117
|
+
|
|
118
|
+
calc(ast) # => 7
|
|
119
|
+
```
|
|
120
|
+
|
|
121
|
+
### Writing handlers
|
|
122
|
+
|
|
123
|
+
When a helper doesn't fit, a handler is just a proc:
|
|
124
|
+
|
|
125
|
+
- A **prefix** handler gets `(token, parser, table)`.
|
|
126
|
+
- An **infix** handler gets `(left, token, parser, table)`.
|
|
127
|
+
- An infix entry is `[binding_power, handler]`.
|
|
128
|
+
|
|
129
|
+
Inside a handler, `parser.parse(table, min_bp)` parses a sub-expression,
|
|
130
|
+
`parser.expect(:type)` consumes a token or raises, and `parser.peek` and
|
|
131
|
+
`parser.advance` look at and consume tokens. `parser.error!(token, message)` raises a
|
|
132
|
+
`Lilt::ParseError` with the token's position.
|
|
133
|
+
|
|
134
|
+
Procs ignore arguments they don't ask for, so simple handlers stay short.
|
|
135
|
+
|
|
136
|
+
### Application by juxtaposition
|
|
137
|
+
|
|
138
|
+
For functional languages, `f x y` means `(f x) y`. Add a `juxtapose:` entry and any
|
|
139
|
+
token that can *start* an expression, appearing where an operator could go, is treated
|
|
140
|
+
as an invisible, left-associative application operator:
|
|
141
|
+
|
|
142
|
+
```ruby
|
|
143
|
+
Var = Data.define(:name)
|
|
144
|
+
App = Data.define(:fn, :arg)
|
|
145
|
+
|
|
146
|
+
LC = {
|
|
147
|
+
prefix: {
|
|
148
|
+
ident: proc { |tok| Var.new(tok.value.to_sym) },
|
|
149
|
+
lparen: Lilt.group(:rparen),
|
|
150
|
+
},
|
|
151
|
+
infix: {
|
|
152
|
+
plus: Lilt.binary(10) { |l, r| Add.new(l, r) },
|
|
153
|
+
},
|
|
154
|
+
juxtapose: [100, proc { |fn, arg| App.new(fn, arg) }],
|
|
155
|
+
}
|
|
156
|
+
|
|
157
|
+
Lilt::Sexp.print(Lilt.parse(LC, Lilt::Lexer.lex("f x y + g y", RULES)))
|
|
158
|
+
# => "(add (app (app (var f) (var x)) (var y)) (app (var g) (var y)))"
|
|
159
|
+
```
|
|
160
|
+
|
|
161
|
+
A real infix operator always wins over juxtaposition, and tokens with no prefix handler
|
|
162
|
+
(like `)`) end an application.
|
|
163
|
+
|
|
164
|
+
### Several tables
|
|
165
|
+
|
|
166
|
+
A grammar can have more than one table (terms and types, say), and a handler switches
|
|
167
|
+
between them just by parsing with a different one. For example, a lambda handler that
|
|
168
|
+
parses `λx:T. body` does `parser.parse(TYPES)` for the annotation and
|
|
169
|
+
`parser.parse(TERMS)` for the body. See `examples/stlc.rb`.
|
|
170
|
+
|
|
171
|
+
You need a separate table for each syntactic category: a part of the grammar where the
|
|
172
|
+
same tokens mean different things, or different operators apply. Many small languages,
|
|
173
|
+
like the untyped lambda calculus, have just one category, so one table is all they need.
|
|
174
|
+
|
|
175
|
+
### Errors and positions
|
|
176
|
+
|
|
177
|
+
Parse errors say where they happened:
|
|
178
|
+
|
|
179
|
+
```ruby
|
|
180
|
+
Lilt.parse(ARITH, Lilt::Lexer.lex("1 +", RULES))
|
|
181
|
+
# raises Lilt::ParseError: unexpected eof at 1:4
|
|
182
|
+
```
|
|
183
|
+
|
|
184
|
+
For errors found *after* parsing (type errors, say), `Lilt.parse_located` also
|
|
185
|
+
returns a table from each node to the token where it starts. The table is compared by
|
|
186
|
+
identity, so two equal nodes at different places in the source are told apart:
|
|
187
|
+
|
|
188
|
+
```ruby
|
|
189
|
+
ast, positions = Lilt.parse_located(ARITH, Lilt::Lexer.lex("1 + 2", RULES))
|
|
190
|
+
positions[ast.right] # => #<data Lilt::Token type=:int, value="2", line=1, col=5>
|
|
191
|
+
```
|
|
192
|
+
|
|
193
|
+
This keeps positions out of the AST, so structural equality still works. The STLC
|
|
194
|
+
example uses this to report type errors with a line and column, without its type checker
|
|
195
|
+
knowing anything about positions.
|
|
196
|
+
|
|
197
|
+
## S-expressions
|
|
198
|
+
|
|
199
|
+
`Lilt::Sexp.read` turns s-expression text into nested arrays of symbols and integers,
|
|
200
|
+
and `Lilt::Sexp.print` goes the other way. It also prints any `Data` node, using the
|
|
201
|
+
snake-cased class name as the tag:
|
|
202
|
+
|
|
203
|
+
```ruby
|
|
204
|
+
Lilt::Sexp.read("(lambda (x Bool) (f x 42))")
|
|
205
|
+
# => [:lambda, [:x, :Bool], [:f, :x, 42]]
|
|
206
|
+
```
|
|
207
|
+
|
|
208
|
+
This is handy for prototyping a language before its syntax has settled: write a small
|
|
209
|
+
function that turns the arrays into your AST, and get to the semantics without writing
|
|
210
|
+
a grammar at all. Later, the Pratt front end can produce the same AST.
|
|
211
|
+
|
|
212
|
+
## What is Pratt parsing?
|
|
213
|
+
|
|
214
|
+
The core loop is the one from Vaughan Pratt's 1973 paper *Top Down Operator
|
|
215
|
+
Precedence*, as popularized by Douglas Crockford. Prefix handlers are Pratt's *nud*s,
|
|
216
|
+
infix handlers are his *led*s, and parsing continues while the next token's binding power
|
|
217
|
+
is greater than the current minimum.
|
|
218
|
+
|
|
219
|
+
Lilt adds two things that aren't in the original:
|
|
220
|
+
|
|
221
|
+
- **Juxtaposition**, for function application.
|
|
222
|
+
- **Multiple tables** per grammar.
|
|
223
|
+
|
|
224
|
+
Within any one table, it's the standard algorithm. The lexer and the s-expression reader
|
|
225
|
+
aren't Pratt parsing; the reader is ordinary recursive descent.
|
|
226
|
+
|
|
227
|
+
## The STLC example
|
|
228
|
+
|
|
229
|
+
[`examples/stlc.rb`](examples/stlc.rb) is a complete simply typed lambda calculus with
|
|
230
|
+
booleans, built on Lilt in about 200 lines. It has:
|
|
231
|
+
|
|
232
|
+
- a Pratt grammar with two tables, terms and types
|
|
233
|
+
- an s-expression front end that produces the same AST, and a test checking that the two agree
|
|
234
|
+
- a type checker whose errors report line and column
|
|
235
|
+
- an environment-based evaluator with closures
|
|
236
|
+
- a De Bruijn index pass, with an evaluator for nameless terms
|
|
237
|
+
|
|
238
|
+
```ruby
|
|
239
|
+
STLC.interpret("(λx:Bool. if x then false else true) true")
|
|
240
|
+
# => #<data STLC::Bool value=false>
|
|
241
|
+
|
|
242
|
+
STLC.interpret("true false")
|
|
243
|
+
# raises STLC::TypeError: expected a function, got (t_bool) at 1:1
|
|
244
|
+
```
|
|
245
|
+
|
|
246
|
+
From a clone, try it with:
|
|
247
|
+
|
|
248
|
+
```sh
|
|
249
|
+
ruby -Ilib -r./examples/stlc -e 'p STLC.interpret("(λx:Bool. x) true")'
|
|
250
|
+
```
|
|
251
|
+
|
|
252
|
+
Its tests are in [`examples/stlc_test.rb`](examples/stlc_test.rb).
|
|
253
|
+
|
|
254
|
+
## Development
|
|
255
|
+
|
|
256
|
+
```sh
|
|
257
|
+
bundle install
|
|
258
|
+
bundle exec rake
|
|
259
|
+
```
|
|
260
|
+
|
|
261
|
+
`rake` runs the tests for the gem and the examples.
|
|
262
|
+
|
|
263
|
+
## License
|
|
264
|
+
|
|
265
|
+
MIT. See [LICENSE.txt](LICENSE.txt).
|
data/lib/lilt/lexer.rb
ADDED
|
@@ -0,0 +1,42 @@
|
|
|
1
|
+
# frozen_string_literal: true
|
|
2
|
+
|
|
3
|
+
require "strscan"
|
|
4
|
+
|
|
5
|
+
module Lilt
|
|
6
|
+
Token = Data.define(:type, :value, :line, :col)
|
|
7
|
+
|
|
8
|
+
class LexError < StandardError; end
|
|
9
|
+
|
|
10
|
+
module Lexer
|
|
11
|
+
extend self
|
|
12
|
+
|
|
13
|
+
def lex(source, rules, keywords: [])
|
|
14
|
+
scanner = StringScanner.new(source)
|
|
15
|
+
tokens = []
|
|
16
|
+
line, line_start = 1, 0
|
|
17
|
+
|
|
18
|
+
until scanner.eos?
|
|
19
|
+
start = scanner.charpos
|
|
20
|
+
col = start - line_start + 1
|
|
21
|
+
rule = rules.find { |pattern, _| scanner.scan(pattern) }
|
|
22
|
+
raise LexError, "unexpected '#{source[start]}' at #{line}:#{col}" unless rule
|
|
23
|
+
|
|
24
|
+
text = scanner.matched
|
|
25
|
+
tokens << Token.new(token_type(rule.last, text, keywords), text, line, col) if rule.last
|
|
26
|
+
line, line_start = advance(line, line_start, start, text)
|
|
27
|
+
end
|
|
28
|
+
|
|
29
|
+
tokens << Token.new(:eof, nil, line, scanner.charpos - line_start + 1)
|
|
30
|
+
end
|
|
31
|
+
|
|
32
|
+
private
|
|
33
|
+
|
|
34
|
+
def token_type(type, text, keywords) = keywords.include?(text) ? text.to_sym : type
|
|
35
|
+
|
|
36
|
+
# Returns the [line, line_start] in effect after consuming +text+ from +start+.
|
|
37
|
+
def advance(line, line_start, start, text)
|
|
38
|
+
nl = text.rindex("\n") or return [line, line_start]
|
|
39
|
+
[line + text.count("\n"), start + nl + 1]
|
|
40
|
+
end
|
|
41
|
+
end
|
|
42
|
+
end
|
data/lib/lilt/parser.rb
ADDED
|
@@ -0,0 +1,104 @@
|
|
|
1
|
+
# frozen_string_literal: true
|
|
2
|
+
|
|
3
|
+
module Lilt
|
|
4
|
+
extend self
|
|
5
|
+
|
|
6
|
+
class ParseError < StandardError; end
|
|
7
|
+
|
|
8
|
+
def parse(table, tokens) = parse_located(table, tokens).first
|
|
9
|
+
|
|
10
|
+
# Like parse, but also returns a Hash (compared by identity) from each node
|
|
11
|
+
# a handler built to the token where that node's expression starts.
|
|
12
|
+
def parse_located(table, tokens)
|
|
13
|
+
parser = Parser.new(tokens)
|
|
14
|
+
ast = parser.parse(table).tap { parser.expect(:eof) }
|
|
15
|
+
[ast, parser.positions]
|
|
16
|
+
end
|
|
17
|
+
|
|
18
|
+
# Builds an infix entry for a binary operator. The block receives the
|
|
19
|
+
# left and right operands (and the operator token) and returns the node.
|
|
20
|
+
def binary(bp, assoc = :left, &build)
|
|
21
|
+
right_bp = assoc == :right ? bp - 1 : bp
|
|
22
|
+
[bp, proc { |left, tok, p, table| build.call(left, p.parse(table, right_bp), tok) }]
|
|
23
|
+
end
|
|
24
|
+
|
|
25
|
+
# Builds a prefix handler for a prefix operator. The operand is parsed at
|
|
26
|
+
# +bp+, so it takes in exactly the infix operators that bind tighter than
|
|
27
|
+
# +bp+. The block receives the operand (and the operator token).
|
|
28
|
+
def prefix(bp, &build)
|
|
29
|
+
proc { |tok, p, table| build.call(p.parse(table, bp), tok) }
|
|
30
|
+
end
|
|
31
|
+
|
|
32
|
+
# Builds a prefix handler for a bracketing token: parses one expression in
|
|
33
|
+
# the current table, then expects +close+. Returns the inner expression.
|
|
34
|
+
def group(close)
|
|
35
|
+
proc { |_tok, p, table| p.parse(table).tap { p.expect(close) } }
|
|
36
|
+
end
|
|
37
|
+
|
|
38
|
+
# A cursor over a token stream, which always ends with an :eof token.
|
|
39
|
+
# Invariant: @pos always indexes a token; advance stops at :eof.
|
|
40
|
+
class Parser
|
|
41
|
+
attr_reader :positions
|
|
42
|
+
|
|
43
|
+
def initialize(tokens)
|
|
44
|
+
@tokens = tokens
|
|
45
|
+
@pos = 0
|
|
46
|
+
@positions = {}.compare_by_identity
|
|
47
|
+
end
|
|
48
|
+
|
|
49
|
+
def parse(table, min_bp = 0)
|
|
50
|
+
start = advance
|
|
51
|
+
nud = table.fetch(:prefix)[start.type] or error!(start, "unexpected #{start.type}")
|
|
52
|
+
left = locate(nud.call(start, self, table), start)
|
|
53
|
+
|
|
54
|
+
loop do
|
|
55
|
+
bp, led = led_for(table, peek)
|
|
56
|
+
break if bp.nil? || bp <= min_bp
|
|
57
|
+
left = locate(led.call(left), start)
|
|
58
|
+
end
|
|
59
|
+
|
|
60
|
+
left
|
|
61
|
+
end
|
|
62
|
+
|
|
63
|
+
def peek = @tokens[@pos]
|
|
64
|
+
|
|
65
|
+
def expect(type)
|
|
66
|
+
tok = peek
|
|
67
|
+
error!(tok, "expected #{type}, got #{tok.type}") unless tok.type == type
|
|
68
|
+
advance
|
|
69
|
+
end
|
|
70
|
+
|
|
71
|
+
def advance
|
|
72
|
+
tok = @tokens[@pos]
|
|
73
|
+
@pos += 1 unless tok.type == :eof
|
|
74
|
+
tok
|
|
75
|
+
end
|
|
76
|
+
|
|
77
|
+
def error!(tok, message) = raise(ParseError, "#{message} at #{tok.line}:#{tok.col}")
|
|
78
|
+
|
|
79
|
+
private
|
|
80
|
+
|
|
81
|
+
# Records where +node+ starts, unless an inner parse already did (as when
|
|
82
|
+
# a group handler returns its inner expression). Returns +node+.
|
|
83
|
+
def locate(node, tok)
|
|
84
|
+
@positions[node] ||= tok
|
|
85
|
+
node
|
|
86
|
+
end
|
|
87
|
+
|
|
88
|
+
# If +tok+ can continue an expression, returns its binding power and a
|
|
89
|
+
# callable that extends +left+. An infix entry consumes the operator
|
|
90
|
+
# token; juxtaposition applies when +tok+ can instead *start* an
|
|
91
|
+
# expression, and consumes nothing before parsing the right side.
|
|
92
|
+
def led_for(table, tok)
|
|
93
|
+
if (infix = table.fetch(:infix, {})[tok.type])
|
|
94
|
+
bp, handler = infix
|
|
95
|
+
return [bp, ->(left) { handler.call(left, advance, self, table) }]
|
|
96
|
+
end
|
|
97
|
+
|
|
98
|
+
bp, build = table[:juxtapose]
|
|
99
|
+
return unless bp && table.fetch(:prefix).key?(tok.type)
|
|
100
|
+
|
|
101
|
+
[bp, ->(left) { build.call(left, parse(table, bp)) }]
|
|
102
|
+
end
|
|
103
|
+
end
|
|
104
|
+
end
|
data/lib/lilt/sexp.rb
ADDED
|
@@ -0,0 +1,54 @@
|
|
|
1
|
+
# frozen_string_literal: true
|
|
2
|
+
|
|
3
|
+
require_relative "lexer"
|
|
4
|
+
require_relative "parser"
|
|
5
|
+
|
|
6
|
+
module Lilt
|
|
7
|
+
module Sexp
|
|
8
|
+
extend self
|
|
9
|
+
|
|
10
|
+
RULES = [[/\s+/, nil], ["(", :lparen], [")", :rparen], [/[^\s()]+/, :atom]]
|
|
11
|
+
|
|
12
|
+
def read(source)
|
|
13
|
+
parser = Parser.new(Lexer.lex(source, RULES))
|
|
14
|
+
datum(parser).tap { parser.expect(:eof) }
|
|
15
|
+
end
|
|
16
|
+
|
|
17
|
+
def print(node)
|
|
18
|
+
case node
|
|
19
|
+
in Data then list([tag(node), *node.deconstruct])
|
|
20
|
+
in Array then list(node)
|
|
21
|
+
else node.to_s
|
|
22
|
+
end
|
|
23
|
+
end
|
|
24
|
+
|
|
25
|
+
private
|
|
26
|
+
|
|
27
|
+
def list(items) = "(#{items.map { print(_1) }.join(" ")})"
|
|
28
|
+
|
|
29
|
+
def datum(parser)
|
|
30
|
+
tok = parser.advance
|
|
31
|
+
case tok.type
|
|
32
|
+
when :lparen then items(parser)
|
|
33
|
+
when :atom then atom(tok.value)
|
|
34
|
+
else parser.error!(tok, "unexpected #{tok.type}")
|
|
35
|
+
end
|
|
36
|
+
end
|
|
37
|
+
|
|
38
|
+
def atom(text) = text.match?(/\A-?\d+\z/) ? text.to_i : text.to_sym
|
|
39
|
+
|
|
40
|
+
def items(parser)
|
|
41
|
+
list = []
|
|
42
|
+
list << datum(parser) until parser.peek.type == :rparen
|
|
43
|
+
parser.expect(:rparen)
|
|
44
|
+
list
|
|
45
|
+
end
|
|
46
|
+
|
|
47
|
+
def tag(node)
|
|
48
|
+
node.class.name.split("::").last
|
|
49
|
+
.gsub(/([A-Z]+)([A-Z][a-z])/, '\1_\2')
|
|
50
|
+
.gsub(/([a-z\d])([A-Z])/, '\1_\2')
|
|
51
|
+
.downcase
|
|
52
|
+
end
|
|
53
|
+
end
|
|
54
|
+
end
|
data/lib/lilt/version.rb
ADDED
data/lib/lilt.rb
ADDED
metadata
ADDED
|
@@ -0,0 +1,52 @@
|
|
|
1
|
+
--- !ruby/object:Gem::Specification
|
|
2
|
+
name: lilt
|
|
3
|
+
version: !ruby/object:Gem::Version
|
|
4
|
+
version: 0.1.0
|
|
5
|
+
platform: ruby
|
|
6
|
+
authors:
|
|
7
|
+
- Phil Crissman
|
|
8
|
+
bindir: bin
|
|
9
|
+
cert_chain: []
|
|
10
|
+
date: 1980-01-02 00:00:00.000000000 Z
|
|
11
|
+
dependencies: []
|
|
12
|
+
description: Lilt provides a table-driven lexer, a Pratt (top-down operator precedence)
|
|
13
|
+
parser engine extended with juxtaposition and multiple expression tables, position
|
|
14
|
+
tracking, and an s-expression reader and printer. Grammars are plain data; ASTs
|
|
15
|
+
are whatever your handlers build.
|
|
16
|
+
email:
|
|
17
|
+
- phil.crissman@gmail.com
|
|
18
|
+
executables: []
|
|
19
|
+
extensions: []
|
|
20
|
+
extra_rdoc_files: []
|
|
21
|
+
files:
|
|
22
|
+
- LICENSE.txt
|
|
23
|
+
- README.md
|
|
24
|
+
- lib/lilt.rb
|
|
25
|
+
- lib/lilt/lexer.rb
|
|
26
|
+
- lib/lilt/parser.rb
|
|
27
|
+
- lib/lilt/sexp.rb
|
|
28
|
+
- lib/lilt/version.rb
|
|
29
|
+
homepage: https://github.com/philcrissman/lilt
|
|
30
|
+
licenses:
|
|
31
|
+
- MIT
|
|
32
|
+
metadata:
|
|
33
|
+
homepage_uri: https://github.com/philcrissman/lilt
|
|
34
|
+
source_code_uri: https://github.com/philcrissman/lilt
|
|
35
|
+
rdoc_options: []
|
|
36
|
+
require_paths:
|
|
37
|
+
- lib
|
|
38
|
+
required_ruby_version: !ruby/object:Gem::Requirement
|
|
39
|
+
requirements:
|
|
40
|
+
- - ">="
|
|
41
|
+
- !ruby/object:Gem::Version
|
|
42
|
+
version: 3.2.0
|
|
43
|
+
required_rubygems_version: !ruby/object:Gem::Requirement
|
|
44
|
+
requirements:
|
|
45
|
+
- - ">="
|
|
46
|
+
- !ruby/object:Gem::Version
|
|
47
|
+
version: '0'
|
|
48
|
+
requirements: []
|
|
49
|
+
rubygems_version: 3.6.7
|
|
50
|
+
specification_version: 4
|
|
51
|
+
summary: A small toolkit for building lexers and Pratt parsers for little languages.
|
|
52
|
+
test_files: []
|