lilt 0.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
checksums.yaml ADDED
@@ -0,0 +1,7 @@
1
+ ---
2
+ SHA256:
3
+ metadata.gz: 27310c339a7875bd8f1e6705459689c05dc978c0c5dc6e1d59534b650493528e
4
+ data.tar.gz: 34cb476a1f8e961d55af835aeb5e490b2fbb7dde94d3b58fe3647c90181b40b2
5
+ SHA512:
6
+ metadata.gz: 7159ca1d750e56ee6c0bbb8c958db5a22188dd8784895a8c9765a5f9038ec8612258914ca0cd6786e003d1f7883f99c3d0de020aa72fc7594f9fdf3324ca0ed4
7
+ data.tar.gz: 4fbb86254219e96138bf7d8409067791dcf3e4d5a43e0e1bc3b595969cfc263b6f2e60ff7a9aff9947cbee5e508dba8311ea44bc632eba5137fc80556daac37a
data/LICENSE.txt ADDED
@@ -0,0 +1,21 @@
1
+ The MIT License (MIT)
2
+
3
+ Copyright (c) 2026 Phil Crissman
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in
13
+ all copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN
21
+ THE SOFTWARE.
data/README.md ADDED
@@ -0,0 +1,265 @@
1
+ # Lilt
2
+
3
+ A **Li**ttle **L**anguage **T**oolkit.
4
+
5
+ I've been writing a lot of small languages, mostly lambda calculi and similar experiments,
6
+ from [TAPL](https://www.cis.upenn.edu/~bcpierce/tapl/) or other sources. I started to notice that the lexer and parser were all generally
7
+ the same, and were not the most interesting part of the process, so I thought I would try
8
+ to create a library to make it easier to get a lexer and parser going for small languages.
9
+
10
+ It has:
11
+
12
+ - **A table-driven lexer.** You write the rules for recognizing tokens, and get back the
13
+ list of tokens with line and column numbers.
14
+ - **A Pratt parser engine.** You write tables of handlers, and get back whatever AST
15
+ your handlers build.
16
+ - **An s-expression reader and printer.** It gives you a second front end for free, and
17
+ a readable way to print any AST.
18
+
19
+ Grammars are plain data (arrays and hashes of procs), and the AST is yours: Lilt
20
+ doesn't define node types. I've been using `Data.define` for nodes, which gives you
21
+ structural equality and pattern matching for free.
22
+
23
+ ## Installation
24
+
25
+ Add it to your Gemfile:
26
+
27
+ ```ruby
28
+ gem "lilt"
29
+ ```
30
+
31
+ or install it directly:
32
+
33
+ ```sh
34
+ gem install lilt
35
+ ```
36
+
37
+ It requires Ruby 3.2 or newer (for `Data`).
38
+
39
+ ## Lexing
40
+
41
+ A lexer is a list of rules, each a pattern and a token type. Rules are tried in order,
42
+ and the first match wins. A string matches literally; a regex matches a pattern; a
43
+ `nil` type means "skip this".
44
+
45
+ ```ruby
46
+ require "lilt"
47
+
48
+ RULES = [
49
+ [/\s+/, nil], # skip whitespace
50
+ ["+", :plus],
51
+ ["*", :star],
52
+ ["(", :lparen],
53
+ [")", :rparen],
54
+ [/\d+/, :int],
55
+ [/[a-z]\w*/, :ident],
56
+ ]
57
+
58
+ Lilt::Lexer.lex("let x", RULES, keywords: %w[let])
59
+ # => [#<data Lilt::Token type=:let, value="let", line=1, col=1>,
60
+ # #<data Lilt::Token type=:ident, value="x", line=1, col=5>,
61
+ # #<data Lilt::Token type=:eof, value=nil, line=1, col=6>]
62
+ ```
63
+
64
+ `keywords:` promotes any token whose whole text is a keyword to its own type, so `let`
65
+ becomes `:let` while `letter` stays an `:ident`. Every token list ends with `:eof`.
66
+ Columns count characters, not bytes, so `λ` is fine. Input that no rule matches raises
67
+ `Lilt::LexError` with its position.
68
+
69
+ ## Parsing
70
+
71
+ A grammar table has `prefix:` handlers (for tokens that start an expression) and
72
+ `infix:` entries (for tokens that continue one). An infix entry is a binding power and a
73
+ handler; `Lilt.binary` builds the common case for you.
74
+
75
+ ```ruby
76
+ Num = Data.define(:value)
77
+ Add = Data.define(:left, :right)
78
+ Mul = Data.define(:left, :right)
79
+
80
+ ARITH = {
81
+ prefix: {
82
+ int: proc { |tok| Num.new(tok.value.to_i) },
83
+ lparen: Lilt.group(:rparen),
84
+ },
85
+ infix: {
86
+ plus: Lilt.binary(10) { |l, r| Add.new(l, r) },
87
+ star: Lilt.binary(20) { |l, r| Mul.new(l, r) },
88
+ },
89
+ }
90
+
91
+ ast = Lilt.parse(ARITH, Lilt::Lexer.lex("1 + 2 * 3", RULES))
92
+ # => #<data Add left=#<data Num value=1>, right=#<data Mul left=#<data Num value=2>, right=#<data Num value=3>>>
93
+
94
+ Lilt::Sexp.print(ast)
95
+ # => "(add (num 1) (mul (num 2) (num 3)))"
96
+ ```
97
+
98
+ A higher binding power binds more tightly. `Lilt.binary(bp)` is left-associative;
99
+ `Lilt.binary(bp, :right)` is right-associative. `Lilt.group(:rparen)` parses one
100
+ expression and then expects the closing token.
101
+
102
+ For prefix operators like unary minus, `Lilt.prefix(bp) { |operand| Neg.new(operand) }`
103
+ parses its operand at `bp`. With a high `bp` (say 70), `-a * b` means `(-a) * b`; with
104
+ one between `*` and `^`, `-a ^ b` means `-(a ^ b)`. A token can have both a prefix and
105
+ an infix handler, so `-` can mean negation and subtraction.
106
+
107
+ Since the AST is just data, doing something with it is a `case`/`in`:
108
+
109
+ ```ruby
110
+ def calc(e)
111
+ case e
112
+ in Num[n] then n
113
+ in Add[l, r] then calc(l) + calc(r)
114
+ in Mul[l, r] then calc(l) * calc(r)
115
+ end
116
+ end
117
+
118
+ calc(ast) # => 7
119
+ ```
120
+
121
+ ### Writing handlers
122
+
123
+ When a helper doesn't fit, a handler is just a proc:
124
+
125
+ - A **prefix** handler gets `(token, parser, table)`.
126
+ - An **infix** handler gets `(left, token, parser, table)`.
127
+ - An infix entry is `[binding_power, handler]`.
128
+
129
+ Inside a handler, `parser.parse(table, min_bp)` parses a sub-expression,
130
+ `parser.expect(:type)` consumes a token or raises, and `parser.peek` and
131
+ `parser.advance` look at and consume tokens. `parser.error!(token, message)` raises a
132
+ `Lilt::ParseError` with the token's position.
133
+
134
+ Procs ignore arguments they don't ask for, so simple handlers stay short.
135
+
136
+ ### Application by juxtaposition
137
+
138
+ For functional languages, `f x y` means `(f x) y`. Add a `juxtapose:` entry and any
139
+ token that can *start* an expression, appearing where an operator could go, is treated
140
+ as an invisible, left-associative application operator:
141
+
142
+ ```ruby
143
+ Var = Data.define(:name)
144
+ App = Data.define(:fn, :arg)
145
+
146
+ LC = {
147
+ prefix: {
148
+ ident: proc { |tok| Var.new(tok.value.to_sym) },
149
+ lparen: Lilt.group(:rparen),
150
+ },
151
+ infix: {
152
+ plus: Lilt.binary(10) { |l, r| Add.new(l, r) },
153
+ },
154
+ juxtapose: [100, proc { |fn, arg| App.new(fn, arg) }],
155
+ }
156
+
157
+ Lilt::Sexp.print(Lilt.parse(LC, Lilt::Lexer.lex("f x y + g y", RULES)))
158
+ # => "(add (app (app (var f) (var x)) (var y)) (app (var g) (var y)))"
159
+ ```
160
+
161
+ A real infix operator always wins over juxtaposition, and tokens with no prefix handler
162
+ (like `)`) end an application.
163
+
164
+ ### Several tables
165
+
166
+ A grammar can have more than one table (terms and types, say), and a handler switches
167
+ between them just by parsing with a different one. For example, a lambda handler that
168
+ parses `λx:T. body` does `parser.parse(TYPES)` for the annotation and
169
+ `parser.parse(TERMS)` for the body. See `examples/stlc.rb`.
170
+
171
+ You need a separate table for each syntactic category: a part of the grammar where the
172
+ same tokens mean different things, or different operators apply. Many small languages,
173
+ like the untyped lambda calculus, have just one category, so one table is all they need.
174
+
175
+ ### Errors and positions
176
+
177
+ Parse errors say where they happened:
178
+
179
+ ```ruby
180
+ Lilt.parse(ARITH, Lilt::Lexer.lex("1 +", RULES))
181
+ # raises Lilt::ParseError: unexpected eof at 1:4
182
+ ```
183
+
184
+ For errors found *after* parsing (type errors, say), `Lilt.parse_located` also
185
+ returns a table from each node to the token where it starts. The table is compared by
186
+ identity, so two equal nodes at different places in the source are told apart:
187
+
188
+ ```ruby
189
+ ast, positions = Lilt.parse_located(ARITH, Lilt::Lexer.lex("1 + 2", RULES))
190
+ positions[ast.right] # => #<data Lilt::Token type=:int, value="2", line=1, col=5>
191
+ ```
192
+
193
+ This keeps positions out of the AST, so structural equality still works. The STLC
194
+ example uses this to report type errors with a line and column, without its type checker
195
+ knowing anything about positions.
196
+
197
+ ## S-expressions
198
+
199
+ `Lilt::Sexp.read` turns s-expression text into nested arrays of symbols and integers,
200
+ and `Lilt::Sexp.print` goes the other way. It also prints any `Data` node, using the
201
+ snake-cased class name as the tag:
202
+
203
+ ```ruby
204
+ Lilt::Sexp.read("(lambda (x Bool) (f x 42))")
205
+ # => [:lambda, [:x, :Bool], [:f, :x, 42]]
206
+ ```
207
+
208
+ This is handy for prototyping a language before its syntax has settled: write a small
209
+ function that turns the arrays into your AST, and get to the semantics without writing
210
+ a grammar at all. Later, the Pratt front end can produce the same AST.
211
+
212
+ ## What is Pratt parsing?
213
+
214
+ The core loop is the one from Vaughan Pratt's 1973 paper *Top Down Operator
215
+ Precedence*, as popularized by Douglas Crockford. Prefix handlers are Pratt's *nud*s,
216
+ infix handlers are his *led*s, and parsing continues while the next token's binding power
217
+ is greater than the current minimum.
218
+
219
+ Lilt adds two things that aren't in the original:
220
+
221
+ - **Juxtaposition**, for function application.
222
+ - **Multiple tables** per grammar.
223
+
224
+ Within any one table, it's the standard algorithm. The lexer and the s-expression reader
225
+ aren't Pratt parsing; the reader is ordinary recursive descent.
226
+
227
+ ## The STLC example
228
+
229
+ [`examples/stlc.rb`](examples/stlc.rb) is a complete simply typed lambda calculus with
230
+ booleans, built on Lilt in about 200 lines. It has:
231
+
232
+ - a Pratt grammar with two tables, terms and types
233
+ - an s-expression front end that produces the same AST, and a test checking that the two agree
234
+ - a type checker whose errors report line and column
235
+ - an environment-based evaluator with closures
236
+ - a De Bruijn index pass, with an evaluator for nameless terms
237
+
238
+ ```ruby
239
+ STLC.interpret("(λx:Bool. if x then false else true) true")
240
+ # => #<data STLC::Bool value=false>
241
+
242
+ STLC.interpret("true false")
243
+ # raises STLC::TypeError: expected a function, got (t_bool) at 1:1
244
+ ```
245
+
246
+ From a clone, try it with:
247
+
248
+ ```sh
249
+ ruby -Ilib -r./examples/stlc -e 'p STLC.interpret("(λx:Bool. x) true")'
250
+ ```
251
+
252
+ Its tests are in [`examples/stlc_test.rb`](examples/stlc_test.rb).
253
+
254
+ ## Development
255
+
256
+ ```sh
257
+ bundle install
258
+ bundle exec rake
259
+ ```
260
+
261
+ `rake` runs the tests for the gem and the examples.
262
+
263
+ ## License
264
+
265
+ MIT. See [LICENSE.txt](LICENSE.txt).
data/lib/lilt/lexer.rb ADDED
@@ -0,0 +1,42 @@
1
+ # frozen_string_literal: true
2
+
3
+ require "strscan"
4
+
5
+ module Lilt
6
+ Token = Data.define(:type, :value, :line, :col)
7
+
8
+ class LexError < StandardError; end
9
+
10
+ module Lexer
11
+ extend self
12
+
13
+ def lex(source, rules, keywords: [])
14
+ scanner = StringScanner.new(source)
15
+ tokens = []
16
+ line, line_start = 1, 0
17
+
18
+ until scanner.eos?
19
+ start = scanner.charpos
20
+ col = start - line_start + 1
21
+ rule = rules.find { |pattern, _| scanner.scan(pattern) }
22
+ raise LexError, "unexpected '#{source[start]}' at #{line}:#{col}" unless rule
23
+
24
+ text = scanner.matched
25
+ tokens << Token.new(token_type(rule.last, text, keywords), text, line, col) if rule.last
26
+ line, line_start = advance(line, line_start, start, text)
27
+ end
28
+
29
+ tokens << Token.new(:eof, nil, line, scanner.charpos - line_start + 1)
30
+ end
31
+
32
+ private
33
+
34
+ def token_type(type, text, keywords) = keywords.include?(text) ? text.to_sym : type
35
+
36
+ # Returns the [line, line_start] in effect after consuming +text+ from +start+.
37
+ def advance(line, line_start, start, text)
38
+ nl = text.rindex("\n") or return [line, line_start]
39
+ [line + text.count("\n"), start + nl + 1]
40
+ end
41
+ end
42
+ end
@@ -0,0 +1,104 @@
1
+ # frozen_string_literal: true
2
+
3
+ module Lilt
4
+ extend self
5
+
6
+ class ParseError < StandardError; end
7
+
8
+ def parse(table, tokens) = parse_located(table, tokens).first
9
+
10
+ # Like parse, but also returns a Hash (compared by identity) from each node
11
+ # a handler built to the token where that node's expression starts.
12
+ def parse_located(table, tokens)
13
+ parser = Parser.new(tokens)
14
+ ast = parser.parse(table).tap { parser.expect(:eof) }
15
+ [ast, parser.positions]
16
+ end
17
+
18
+ # Builds an infix entry for a binary operator. The block receives the
19
+ # left and right operands (and the operator token) and returns the node.
20
+ def binary(bp, assoc = :left, &build)
21
+ right_bp = assoc == :right ? bp - 1 : bp
22
+ [bp, proc { |left, tok, p, table| build.call(left, p.parse(table, right_bp), tok) }]
23
+ end
24
+
25
+ # Builds a prefix handler for a prefix operator. The operand is parsed at
26
+ # +bp+, so it takes in exactly the infix operators that bind tighter than
27
+ # +bp+. The block receives the operand (and the operator token).
28
+ def prefix(bp, &build)
29
+ proc { |tok, p, table| build.call(p.parse(table, bp), tok) }
30
+ end
31
+
32
+ # Builds a prefix handler for a bracketing token: parses one expression in
33
+ # the current table, then expects +close+. Returns the inner expression.
34
+ def group(close)
35
+ proc { |_tok, p, table| p.parse(table).tap { p.expect(close) } }
36
+ end
37
+
38
+ # A cursor over a token stream, which always ends with an :eof token.
39
+ # Invariant: @pos always indexes a token; advance stops at :eof.
40
+ class Parser
41
+ attr_reader :positions
42
+
43
+ def initialize(tokens)
44
+ @tokens = tokens
45
+ @pos = 0
46
+ @positions = {}.compare_by_identity
47
+ end
48
+
49
+ def parse(table, min_bp = 0)
50
+ start = advance
51
+ nud = table.fetch(:prefix)[start.type] or error!(start, "unexpected #{start.type}")
52
+ left = locate(nud.call(start, self, table), start)
53
+
54
+ loop do
55
+ bp, led = led_for(table, peek)
56
+ break if bp.nil? || bp <= min_bp
57
+ left = locate(led.call(left), start)
58
+ end
59
+
60
+ left
61
+ end
62
+
63
+ def peek = @tokens[@pos]
64
+
65
+ def expect(type)
66
+ tok = peek
67
+ error!(tok, "expected #{type}, got #{tok.type}") unless tok.type == type
68
+ advance
69
+ end
70
+
71
+ def advance
72
+ tok = @tokens[@pos]
73
+ @pos += 1 unless tok.type == :eof
74
+ tok
75
+ end
76
+
77
+ def error!(tok, message) = raise(ParseError, "#{message} at #{tok.line}:#{tok.col}")
78
+
79
+ private
80
+
81
+ # Records where +node+ starts, unless an inner parse already did (as when
82
+ # a group handler returns its inner expression). Returns +node+.
83
+ def locate(node, tok)
84
+ @positions[node] ||= tok
85
+ node
86
+ end
87
+
88
+ # If +tok+ can continue an expression, returns its binding power and a
89
+ # callable that extends +left+. An infix entry consumes the operator
90
+ # token; juxtaposition applies when +tok+ can instead *start* an
91
+ # expression, and consumes nothing before parsing the right side.
92
+ def led_for(table, tok)
93
+ if (infix = table.fetch(:infix, {})[tok.type])
94
+ bp, handler = infix
95
+ return [bp, ->(left) { handler.call(left, advance, self, table) }]
96
+ end
97
+
98
+ bp, build = table[:juxtapose]
99
+ return unless bp && table.fetch(:prefix).key?(tok.type)
100
+
101
+ [bp, ->(left) { build.call(left, parse(table, bp)) }]
102
+ end
103
+ end
104
+ end
data/lib/lilt/sexp.rb ADDED
@@ -0,0 +1,54 @@
1
+ # frozen_string_literal: true
2
+
3
+ require_relative "lexer"
4
+ require_relative "parser"
5
+
6
+ module Lilt
7
+ module Sexp
8
+ extend self
9
+
10
+ RULES = [[/\s+/, nil], ["(", :lparen], [")", :rparen], [/[^\s()]+/, :atom]]
11
+
12
+ def read(source)
13
+ parser = Parser.new(Lexer.lex(source, RULES))
14
+ datum(parser).tap { parser.expect(:eof) }
15
+ end
16
+
17
+ def print(node)
18
+ case node
19
+ in Data then list([tag(node), *node.deconstruct])
20
+ in Array then list(node)
21
+ else node.to_s
22
+ end
23
+ end
24
+
25
+ private
26
+
27
+ def list(items) = "(#{items.map { print(_1) }.join(" ")})"
28
+
29
+ def datum(parser)
30
+ tok = parser.advance
31
+ case tok.type
32
+ when :lparen then items(parser)
33
+ when :atom then atom(tok.value)
34
+ else parser.error!(tok, "unexpected #{tok.type}")
35
+ end
36
+ end
37
+
38
+ def atom(text) = text.match?(/\A-?\d+\z/) ? text.to_i : text.to_sym
39
+
40
+ def items(parser)
41
+ list = []
42
+ list << datum(parser) until parser.peek.type == :rparen
43
+ parser.expect(:rparen)
44
+ list
45
+ end
46
+
47
+ def tag(node)
48
+ node.class.name.split("::").last
49
+ .gsub(/([A-Z]+)([A-Z][a-z])/, '\1_\2')
50
+ .gsub(/([a-z\d])([A-Z])/, '\1_\2')
51
+ .downcase
52
+ end
53
+ end
54
+ end
@@ -0,0 +1,5 @@
1
+ # frozen_string_literal: true
2
+
3
+ module Lilt
4
+ VERSION = "0.1.0"
5
+ end
data/lib/lilt.rb ADDED
@@ -0,0 +1,6 @@
1
+ # frozen_string_literal: true
2
+
3
+ require_relative "lilt/version"
4
+ require_relative "lilt/lexer"
5
+ require_relative "lilt/parser"
6
+ require_relative "lilt/sexp"
metadata ADDED
@@ -0,0 +1,52 @@
1
+ --- !ruby/object:Gem::Specification
2
+ name: lilt
3
+ version: !ruby/object:Gem::Version
4
+ version: 0.1.0
5
+ platform: ruby
6
+ authors:
7
+ - Phil Crissman
8
+ bindir: bin
9
+ cert_chain: []
10
+ date: 1980-01-02 00:00:00.000000000 Z
11
+ dependencies: []
12
+ description: Lilt provides a table-driven lexer, a Pratt (top-down operator precedence)
13
+ parser engine extended with juxtaposition and multiple expression tables, position
14
+ tracking, and an s-expression reader and printer. Grammars are plain data; ASTs
15
+ are whatever your handlers build.
16
+ email:
17
+ - phil.crissman@gmail.com
18
+ executables: []
19
+ extensions: []
20
+ extra_rdoc_files: []
21
+ files:
22
+ - LICENSE.txt
23
+ - README.md
24
+ - lib/lilt.rb
25
+ - lib/lilt/lexer.rb
26
+ - lib/lilt/parser.rb
27
+ - lib/lilt/sexp.rb
28
+ - lib/lilt/version.rb
29
+ homepage: https://github.com/philcrissman/lilt
30
+ licenses:
31
+ - MIT
32
+ metadata:
33
+ homepage_uri: https://github.com/philcrissman/lilt
34
+ source_code_uri: https://github.com/philcrissman/lilt
35
+ rdoc_options: []
36
+ require_paths:
37
+ - lib
38
+ required_ruby_version: !ruby/object:Gem::Requirement
39
+ requirements:
40
+ - - ">="
41
+ - !ruby/object:Gem::Version
42
+ version: 3.2.0
43
+ required_rubygems_version: !ruby/object:Gem::Requirement
44
+ requirements:
45
+ - - ">="
46
+ - !ruby/object:Gem::Version
47
+ version: '0'
48
+ requirements: []
49
+ rubygems_version: 3.6.7
50
+ specification_version: 4
51
+ summary: A small toolkit for building lexers and Pratt parsers for little languages.
52
+ test_files: []