@scinorandex/sparse 0.1.1 → 0.1.3

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (65) hide show
  1. package/.github/workflows/ci.yml +41 -0
  2. package/AGENTS.md +31 -11
  3. package/README.md +289 -29
  4. package/dist/cli.js +113 -24
  5. package/dist/cli.js.map +1 -1
  6. package/dist/generator.d.ts +3 -1
  7. package/dist/generator.js +29 -5
  8. package/dist/generator.js.map +1 -1
  9. package/dist/index.d.ts +16 -4
  10. package/dist/index.js +34 -1
  11. package/dist/index.js.map +1 -1
  12. package/dist/meta/common.d.ts +17 -0
  13. package/dist/meta/common.js +66 -1
  14. package/dist/meta/common.js.map +1 -1
  15. package/dist/meta/selfhosted.d.ts +4 -5
  16. package/dist/meta/selfhosted.js +49 -36
  17. package/dist/meta/selfhosted.js.map +1 -1
  18. package/dist/parser.d.ts +55 -35
  19. package/dist/parser.js +141 -18
  20. package/dist/parser.js.map +1 -1
  21. package/dist/table/selfhosted.d.ts +10 -1
  22. package/dist/table/selfhosted.js +32 -12
  23. package/dist/table/selfhosted.js.map +1 -1
  24. package/dist/table/validate.d.ts +17 -0
  25. package/dist/table/validate.js +84 -0
  26. package/dist/table/validate.js.map +1 -0
  27. package/dist/utils/Stack.d.ts +4 -0
  28. package/dist/utils/Stack.js +16 -1
  29. package/dist/utils/Stack.js.map +1 -1
  30. package/dist/utils/errorWindowBuilder.js +11 -11
  31. package/dist/utils/errorWindowBuilder.js.map +1 -1
  32. package/dist/utils/loadFiles.d.ts +8 -0
  33. package/dist/utils/loadFiles.js +47 -0
  34. package/dist/utils/loadFiles.js.map +1 -0
  35. package/dist/utils/reducers.d.ts +9 -0
  36. package/dist/utils/reducers.js +56 -0
  37. package/dist/utils/reducers.js.map +1 -0
  38. package/example/LoLang/example.ts +16 -15
  39. package/example/kleene-test/example.ts +35 -17
  40. package/example/math/example.ts +12 -7
  41. package/example/selfhosted/example.ts +30 -24
  42. package/opencode.json +1 -1
  43. package/package.json +2 -2
  44. package/src/cli.ts +132 -23
  45. package/src/generator.ts +49 -8
  46. package/src/index.ts +55 -3
  47. package/src/meta/common.ts +116 -0
  48. package/src/meta/selfhosted.ts +42 -11
  49. package/src/parser.ts +243 -51
  50. package/src/table/selfhosted.ts +48 -12
  51. package/src/table/validate.ts +143 -0
  52. package/src/utils/Stack.ts +22 -2
  53. package/src/utils/errorWindowBuilder.ts +15 -13
  54. package/src/utils/loadFiles.ts +54 -0
  55. package/src/utils/reducers.ts +93 -0
  56. package/test/cli.test.ts +144 -0
  57. package/test/codegen.test.ts +75 -0
  58. package/test/grammar.test.ts +243 -0
  59. package/test/helpers.ts +59 -0
  60. package/test/lalr.test.ts +110 -0
  61. package/test/parser.test.ts +377 -0
  62. package/test/table.test.ts +144 -0
  63. package/tsconfig.json +1 -1
  64. package/example/LoLang/table2.txt +0 -497
  65. package/example/complicated/grammar.txt +0 -37
@@ -0,0 +1,41 @@
1
+ name: CI
2
+
3
+ on:
4
+ push:
5
+ branches: [main]
6
+ pull_request:
7
+
8
+ jobs:
9
+ verify:
10
+ runs-on: ubuntu-latest
11
+
12
+ steps:
13
+ - uses: actions/checkout@v4
14
+
15
+ - uses: actions/setup-node@v4
16
+ with:
17
+ node-version: 20
18
+ cache: yarn
19
+
20
+ - run: yarn install --frozen-lockfile
21
+
22
+ - name: Typecheck
23
+ run: yarn tc
24
+
25
+ - name: Test
26
+ run: yarn test
27
+
28
+ - name: Build
29
+ run: yarn build
30
+
31
+ - name: Run the examples
32
+ run: |
33
+ npm install -g tsx
34
+ npx tsx example/math/example.ts
35
+ npx tsx example/kleene-test/example.ts
36
+ npx tsx example/LoLang/example.ts
37
+
38
+ - name: The committed tables are up to date
39
+ run: |
40
+ node dist/cli.js --input=example/math/grammar.txt --output=example/math/table.txt --check
41
+ node dist/cli.js --input=example/LoLang/grammar.txt --output=example/LoLang/table.txt --check
package/AGENTS.md CHANGED
@@ -1,14 +1,16 @@
1
1
  # AGENTS.md
2
2
 
3
3
  `@scinorandex/sparse` — LR(1) parser generator library (TypeScript, CJS via tsc, yarn 1).
4
- Single package: `src/` = library + CLI, `example/` = runnable demos. README.md is the user-facing doc.
4
+ Single package: `src/` = library + CLI, `example/` = runnable demos, `test/` = vitest suite.
5
+ README.md is the user-facing doc (grammar syntax, reducer contract, recovery guide, API reference).
5
6
 
6
7
  ## Commands
7
- - Typecheck: `yarn tc` (covers src/ and example/)
8
+ - Typecheck: `yarn tc` (covers src/, example/ and test/)
8
9
  - Build: `yarn build` → `dist/` (CLI is `dist/cli.js`)
10
+ - Test: `yarn test` (vitest; builds `dist/` first because `test/cli.test.ts` runs the CLI)
9
11
  - Run examples: `tsx example/<name>/example.ts` — tsx is global, not a repo dep; `node file.ts` fails (enums)
10
12
  - Table CLI: `node dist/cli.js --input=<grammar> --output=<table>` (build first); output matches committed tables exactly
11
- - `yarn test` / `yarn bench` (vitest) exist but there are no test/bench files — `yarn test` fails with "No test files found". Verify via examples + `yarn tc` instead.
13
+ - `yarn bench` exists but there are no bench files. Verify with `yarn tc && yarn test` plus the examples.
12
14
  - `dist/` and `yarn.lock` are gitignored.
13
15
 
14
16
  ## Self-hosting / codegen (critical, easy to break)
@@ -19,22 +21,40 @@ Single package: `src/` = library + CLI, `example/` = runnable demos. README.md i
19
21
  `02-table-grammar.txt`, run `tsx example/selfhosted/example.ts`, then copy the regenerated
20
22
  `codegen-*.ts` over the two `src/.../states.ts` files. The old (checked-in) parser must still
21
23
  parse the modified grammar (bootstrap constraint).
24
+ - `test/codegen.test.ts` enforces this: it regenerates both state files in memory and compares them to
25
+ the checked-in ones. It compares objects (not file bytes) so the `cp` step stays manual.
26
+ - Changing the *lexer* rules in `grammarLexerGenerator` (e.g. allowing `-` inside identifiers) does NOT
27
+ change the generated states — the token types and positions stay the same.
22
28
 
23
29
  ## Non-obvious behavior
24
- - `?`, `*`, `+` groups and multi-token alternatives are unrolled into extra productions; `*`/`+`
30
+ - `?`, `*`, `+` groups and `|` alternatives are unrolled into extra productions; `*`/`+`
25
31
  create `autogen-N` productions whose reducer is named `autogenerated-kleene` — user reducers must
26
- handle it (see `example/kleene-test/`).
32
+ handle it (see `example/kleene-test/`). `?` on a group that can match nothing produces an empty
33
+ production and is rejected by `validateProductions`.
34
+ - Modifiers only apply to parenthesized groups: `([A] [B])?`, not `[A]?`.
35
+ - `|` alternatives live inside a single token: `[A | B]`, `<A | B>`.
27
36
  - Two construction paths: `Sparse.fromProductions(...)` builds states at runtime (slow, ~30s for
28
- 150 productions) vs `new Sparse({ productions, states: buildStates(table.txt) })` (fast).
37
+ 150 productions) vs `Sparse.fromGrammarFile(...)` / `new Sparse({ productions, states })` (fast).
29
38
  - Reducers are keyed by production name (`<LHS: name>` in the grammar file); nameless productions
30
- have `name: null` (use `reducers[name ?? ""]`). Production index 0 is the accept production —
31
- the parse loop breaks on it without calling the reducer.
39
+ have `name: null`. Production index 0 is the accept production —
40
+ the parse loop breaks on it without calling the reducer. `oldInput.index` is the pre-unrolling index.
32
41
  - Table file format: one line per state, comma-separated `[TERMINAL]=sN|rN` and `<VARIABLE>=M`.
42
+ The reader preserves file order via `getItemsReversed()` (the item lists are built right-recursively).
33
43
  - Errors flow as `Result<T>` (`{ success: false, reason, token }`); `buildProductions` /
34
- `Sparse.fromProductions` are the throwing wrappers.
44
+ `Sparse.fromProductions` / `buildStates` are the throwing wrappers.
45
+ - `GeneratorResult.ActionTable` is sparse: a state whose items only move to other states has no entry.
46
+ `toStates()`/`toTable()` use `Array.from`, not `map`, so state numbering stays aligned; `toTable()`
47
+ throws for such grammars because the file format cannot express an action-less state.
48
+ - LALR(1) mode merges same-core LR(1) states. Same-core states always agree on their shifts, gotos and
49
+ reduce targets, so the conflict branch in `mergeStatesForLALR` is currently unreachable.
35
50
 
36
51
  ## Key files
37
52
  - `src/generator.ts` — LR(1) state construction (first/follow sets, item-set expansion), `toTable()` serialization
38
- - `src/parser.ts` — runtime parser loop + `recover` error-recovery callback API
53
+ - `src/parser.ts` — runtime parser loop + `recover` error-recovery callback API (`insertToken`, `finish`, `crash`)
39
54
  - `src/meta/selfhosted.ts` — grammar-file parser + production unrolling
40
- - `example/` — `math` (prebuilt table), `kleene-test` (`*`/`+`), `LoLang` (large grammar + recovery), `complicated` (grammar only), `selfhosted` (codegen)
55
+ - `src/meta/common.ts` — `Production` type + `validateProductions`
56
+ - `src/table/selfhosted.ts` — table-file parser (`tryBuildStates`), `src/table/validate.ts` — table validation
57
+ - `src/cli.ts` — `--input/--output/--lalr/--check/--stdout/--help`
58
+ - `example/` — `math` (prebuilt table), `kleene-test` (`*`/`+`), `LoLang` (large grammar + recovery, exports
59
+ its lexer for the tests), `selfhosted` (codegen)
60
+ - `test/` — `grammar`, `table`, `parser`, `lalr`, `codegen`, `cli`
package/README.md CHANGED
@@ -1,20 +1,39 @@
1
1
  ## Sparse - Scin's Parsing Library
2
2
 
3
- Sparse allows developers to easily create LR1 parsers and LR1 parsing tables.
3
+ Sparse allows developers to easily create LR(1) parsers and LR(1) parsing tables.
4
4
 
5
5
  **Features:**
6
6
  - Complete Error Handling API - when parsing fails, you have the final say on where parsing stops
7
- - Deferred Reductions - you create the nodes and Sparse builds the tree
7
+ - Deferred Reductions - you create the nodes and Sparse builds the tree
8
+ - Grammar and table validation - mistakes are reported with the line, column and a window into your file
8
9
  - Full [Slex](https://github.com/scinscinscin/slex) integration - define your entire language as a set of DFA and CFG rules
9
10
  - Optional Table Output - use Sparse to build your tables for use in other languages
10
11
 
11
- ## Getting Started
12
+ ## Table of contents
13
+
14
+ - [Sparse - Scin's Parsing Library](#sparse---scins-parsing-library)
15
+ - [Table of contents](#table-of-contents)
16
+ - [Getting started](#getting-started)
17
+ - [Grammar syntax reference](#grammar-syntax-reference)
18
+ - [What Sparse checks for you](#what-sparse-checks-for-you)
19
+ - [Reducers](#reducers)
20
+ - [Repetition with `*` and `+`](#repetition-with--and-)
21
+ - [Error recovery](#error-recovery)
22
+ - [Generating and shipping the parsing table](#generating-and-shipping-the-parsing-table)
23
+ - [API reference](#api-reference)
24
+ - [Reading and writing grammars](#reading-and-writing-grammars)
25
+ - [Building parsers](#building-parsers)
26
+ - [Helpers](#helpers)
27
+ - [Known limitations](#known-limitations)
28
+ - [AI Disclaimer](#ai-disclaimer)
29
+
30
+ ## Getting started
12
31
 
13
32
  1. **Define the CFG of your language.**
14
33
 
15
- The CFG is defined by creating a file containing a list of productions. Variables are identifiers encased in angle brackets like `<STATEMENT>`, while terminals are encased in square-brackets like `[L_COLON]`.
34
+ The CFG is defined by creating a file containing a list of productions. Variables are identifiers encased in angle brackets like `<STATEMENT>`, while terminals are encased in square-brackets like `[L_COLON]`.
16
35
 
17
- An example production includes: `<IF_STATEMENT> : [IF] [L_PAREN] <EXPRESSION> [R_PAREN] <STATEMENT> <ELSE_IF_STATEMENTS> [ELSE] <STATEMENT>;`.
36
+ An example production includes: `<IF_STATEMENT> : [IF] [L_PAREN] <EXPRESSION> [R_PAREN] <STATEMENT> [ELSE] <STATEMENT>;`.
18
37
 
19
38
  An example CFG for a basic MDAS calculator is the following:
20
39
 
@@ -22,16 +41,20 @@ An example CFG for a basic MDAS calculator is the following:
22
41
  <S>: <PROGRAM>;
23
42
  <PROGRAM>: <EXPRESSION> [EOF];
24
43
  <PROGRAM>: [EOF];
44
+
45
+ // Begin parsing arithmetic expressions
25
46
  <EXPRESSION>: <TERM_EXPRESSION>;
26
47
  <TERM_EXPRESSION>: <FACTOR_EXPRESSION> [PLUS] <TERM_EXPRESSION>;
27
48
  <TERM_EXPRESSION>: <FACTOR_EXPRESSION> [MINUS] <TERM_EXPRESSION>;
28
49
  <TERM_EXPRESSION>: <FACTOR_EXPRESSION>;
29
50
  <FACTOR_EXPRESSION>: <ENDPOINT> [STAR] <FACTOR_EXPRESSION>;
30
- <FACTOR_EXPRESSION>: <ENDPOINT> [SLASH] <FACTOR_EXPRESSION>;
51
+ <FACTOR_EXPRESSION>: <ENDPOINT> [FORWARD_SLASH] <FACTOR_EXPRESSION>;
31
52
  <FACTOR_EXPRESSION>: <ENDPOINT>;
32
53
  <ENDPOINT>: [NUMBER];
33
54
  ```
34
55
 
56
+ The first production has to be the "accept" production: a variable that appears nowhere else, whose body is a single variable. It is the only production whose reducer Sparse never calls.
57
+
35
58
  2. **Create your [Slex](https://github.com/scinscinscin/slex) lexer.**
36
59
 
37
60
  ```ts
@@ -61,7 +84,7 @@ lexerGenerator.addRule("number_literal", "${float_number}|${decimal_number}", To
61
84
  const lexer = lexerGenerator.generate(`2.4 + 3.5 * 1 / 456.789`, () => ({}));
62
85
  ```
63
86
 
64
- 3. **Define the node representation.**
87
+ 3. **Define the node representation.**
65
88
 
66
89
  Sparse allows you to build the AST however you want, deferring to your functions when its time to make a reduction, giving you control over the representation.
67
90
 
@@ -71,7 +94,7 @@ type StringifiedNode = (StringifiedNode | string)[];
71
94
  // It doesn't have to be a class. It just has to be a structure that all nodes
72
95
  // in the AST adhere to. Classes allow this to be done easily through subclassing.
73
96
  class Node {
74
- constructor(public readonly nodes: LR1StackSymbol<TokenType, {}, Node>[]) {}
97
+ constructor(public readonly nodes: LR1StackSymbol<TokenType, Metadata, Node>[]) {}
75
98
  toObject(): StringifiedNode {
76
99
  return this.nodes.map((node) => (node.type === "token"
77
100
  ? node.token.lexeme
@@ -81,53 +104,290 @@ class Node {
81
104
  }
82
105
  ```
83
106
 
84
- 4. **Building the productions and the parser.**
107
+ 4. **Building the parser.**
85
108
 
86
- Productions are built using the `buildProductions()` function, which can be passed into Sparse alongside `toStringifiedTokenType`, which converts a numerical TypeScript enum to the name of the terminal.
109
+ With a prebuilt parsing table, `Sparse.fromGrammarFile` reads the grammar, reads the table, and checks that the two belong together:
87
110
 
88
111
  ```ts
112
+ import { Sparse, enumToString } from "@scinorandex/sparse";
113
+
114
+ async function main() {
115
+ const parserGenerator = await Sparse.fromGrammarFile<TokenType, Metadata, Node>({
116
+ grammarPath: "./example/math/grammar.txt",
117
+ tablePath: "./example/math/table.txt",
118
+ toStringifiedTokenType: enumToString<TokenType>(TokenType),
119
+ onWarning: ({ reason, token }) => console.warn(`${token.line}:${token.column} ${reason}`),
120
+ });
121
+
122
+ const parser = parserGenerator.generate(lexer, {
123
+ reducer: (_, { input }) => new Node(input),
124
+ });
125
+
126
+ console.log(parser.parse().result!.toObject());
127
+ }
128
+ ```
129
+
130
+ Generating the table at startup instead (fine for small grammars, slow for big ones):
131
+
132
+ ```ts
133
+ import { buildProductions, Sparse, enumToString } from "@scinorandex/sparse";
134
+
89
135
  async function main() {
90
- const toStringifiedTokenType = (type: TokenType) => TokenType[type];
91
136
  const productions = buildProductions(await fs.readFile("./example/math/grammar.txt", "utf8"));
92
- const parserGenerator = Sparse.fromProductions<TokenType, Metadata, Node>({ productions, toStringifiedTokenType });
137
+ const parserGenerator = Sparse.fromProductions<TokenType, Metadata, Node>({
138
+ productions,
139
+ toStringifiedTokenType: enumToString<TokenType>(TokenType),
140
+ });
93
141
  }
94
142
  ```
95
143
 
96
- 5. **Define your reducer and begin parsing.**
144
+ ## Grammar syntax reference
145
+
146
+ A grammar file is a list of productions. Whitespace is insignificant, `//` starts a line comment, and `/* ... */` spans lines.
147
+
148
+ | Syntax | Meaning |
149
+ | ----------------------------- | ------------------------------------------------------------------------------- |
150
+ | `<A>: <B> [C];` | `A` is made of one `B` followed by one `C` |
151
+ | `<A: name>: ...;` | The production is *named*: its reducer is looked up under `name` |
152
+ | `[TOK: name]` | Names the terminal on the right hand side so it shows up on the reducer's `bag` |
153
+ | `<VAR: name>` | Same, for a variable |
154
+ | `[A | B]` | Alternatives: unrolled into one production per alternative |
155
+ | `<A | B>` | Same, for variables |
156
+ | `([A] [B])?` | Optional group, unrolled into one production with and one without it |
157
+ | `// comment`, `/* comment */` | Comments |
158
+
159
+ Identifiers may contain letters, digits, `_` and `-`, and must start with a letter or `_`.
97
160
 
98
- Finally, you can define the parser's reduction function and optionally define the recovery handling function.
161
+ ### What Sparse checks for you
162
+
163
+ Grammar and table problems are reported as a `Result` failure with the offending token, so the CLI and
164
+ `loadGrammar` can print a window into your file:
165
+
166
+ - a variable used on the right hand side that no production defines
167
+ - an empty right hand side (a production like `<A: a>: ([X])?;` expands to nothing, and Sparse has no support for empty productions)
168
+ - the start symbol appearing on the right hand side
169
+ - two symbols in one production sharing a name (warns: only the last one reaches the `bag`)
170
+ - a production that is never reachable from the start symbol (warns)
171
+
172
+ ```
173
+ $ npx sparse --input=broken.txt --output=table.txt
174
+ Invalid syntax at 2:30: got SEMICOLON (";"), but expected one of [L_ANGLE], [L_BRACKET], [L_PAREN]
175
+
176
+ 2 | <PROGRAM: program>: [NUMBER] (;
177
+ | ~
178
+ 3 |
179
+ ```
180
+
181
+ ## Reducers
182
+
183
+ Reducers are keyed by the name in the grammar, and receive two arguments:
184
+
185
+ ```ts
186
+ type Reducer<TokenType, Metadata, Node> = (
187
+ newInput: { bag: Record<string, Node | Token<TokenType, Metadata>>; name: string | null },
188
+ oldInput: { input: LR1StackSymbol<TokenType, Metadata, Node>[]; index: number },
189
+ ) => Node;
190
+ ```
191
+
192
+ - `newInput.name` is the production's name, or `null` for an unnamed production.
193
+ - `newInput.bag` holds the right hand side symbols that were named in the grammar, keyed by those names. If two symbols share a name, the last one wins, so reach for `oldInput.input` when that happens.
194
+ - `oldInput.input` is the reduced symbols in order, as `{ type: "token", token }` or `{ type: "node", node }`.
195
+ - `oldInput.index` is the production's index *in the grammar file*, before `*`/`?`/`+`/`|` unrolling. Index 0 is the accept production, whose reducer is never called.
196
+
197
+ `defineReducers` turns a map of reducers into the reducer that `generate` wants, and tells you exactly which production has no reducer:
99
198
 
100
199
  ```ts
200
+ import { assertReducersCoverGrammar, defineReducers } from "@scinorandex/sparse";
201
+
202
+ const reducers = {
203
+ program: ({ bag }) => new Node(bag),
204
+ expression: ({ bag }, { input }) => new Node(input),
205
+ };
206
+
207
+ assertReducersCoverGrammar(productions, reducers); // throws if something is missing
208
+
101
209
  const parser = parserGenerator.generate(lexer, {
102
- reducer: (_, { input, index }) => new Node(input),
210
+ reducer: defineReducers(reducers),
103
211
  });
212
+ ```
213
+
214
+ Without the helper the failure looks like this, halfway through a parse:
104
215
 
105
- console.log(parser.parse().result!.toObject());
216
+ ```
217
+ Error while performing the reduction for production 2 ("expression"): I cannot reduce a null bag
106
218
  ```
107
219
 
108
- ## Exporting and Loading the Parsing Table
220
+ ## Repetition with `*` and `+`
109
221
 
110
- It takes a while for Sparse to build states (around 0.5 seconds seconds for a file containing 150 productions). This can be alleviated by generating the parsing table and loading the states directly instead. This method also allows you to edit the parsing table to resolve parsing conflicts.
222
+ `[A]?`, `A B*`, and `A B+` are unrolled into extra productions. For `*` and `+` that means a new
223
+ production named **`autogenerated-kleene`** (with variables called `autogen-0`, `autogen-1`, ...), which
224
+ you **have** to implement: it is handed the items matched so far, and it has to flatten them into a list.
225
+ The first reduction has no `rest`, later ones do:
111
226
 
112
- To create the the states, you can run `npx @scinorandex/sparse --input=<input> --output=<output>`.
227
+ ```ts
228
+ class KleeneNode<T> extends Node {
229
+ contents: T[] = [];
113
230
 
114
- By default, LR(1) states are generated. You can generate LALR(1) states instead by passing the `--lalr` flag (or setting `mode: "lalr1"` in `Sparse.fromProductions`): the resulting table is never larger than the LR(1) one, but generation fails if the grammar is not LALR(1). Using LALR(1) yields a ~35% performance boost over LR(1) for the same grammar (tested on the LoLang example).
231
+ constructor(nodes: LR1StackSymbol<TokenType, Metadata, Node>[], bag: T) {
232
+ super(nodes);
233
+ this.contents.unshift(bag);
234
+ }
115
235
 
116
- Afterwards, you can create a parser with pre-built states like the following:
236
+ add(nodes: LR1StackSymbol<TokenType, Metadata, Node>[], bag: T) {
237
+ this.nodes.unshift(...nodes.slice(0, nodes.length - 1));
238
+ this.contents.unshift(bag);
239
+ return this;
240
+ }
241
+ }
117
242
 
118
- ```ts
119
- async function main() {
120
- const toStringifiedTokenType = (type: TokenType) => TokenType[type];
121
- const productions = buildProductions(await fs.readFile("./example/math/grammar.txt", "utf8"));
243
+ const reducers = {
244
+ "autogenerated-kleene": ({ bag: { rest, ...item } }, { input }) =>
245
+ rest == null ? new KleeneNode(input, item) : rest.add(input, item),
246
+ program: (_, { input }) => new Node(input),
247
+ };
248
+ ```
249
+
250
+ See `example/kleene-test` for the whole thing.
122
251
 
123
- // Notice how the "new Sparse()" constructor is used here instead of "Sparse.fromProductions()"
124
- const states = buildStates(await fs.readFile("./example/math/table.txt", "utf8"));
125
- const parserGenerator = new Sparse<TokenType, Metadata, Node>({ productions, states, toStringifiedTokenType });
252
+ ## Error recovery
253
+
254
+ By default a syntax error throws `LR1ParserGraveError`, which carries the token that could not be parsed:
255
+
256
+ ```ts
257
+ try {
258
+ parser.parse();
259
+ } catch (err) {
260
+ if (err instanceof LR1ParserGraveError) console.error(err.reason, err.currentToken);
126
261
  }
127
262
  ```
128
263
 
264
+ Pass a `recover` function to decide what happens instead. It receives the lexer, both stacks, the table,
265
+ and these helpers:
266
+
267
+ | Helper | What it does |
268
+ | -------------------- | ------------------------------------------------------------------- |
269
+ | `addError(reason)` | Records an error, keeps parsing, and returns it in `parse().errors` |
270
+ | `crash(reason)` | Stops parsing and throws a `LR1ParserGraveError` |
271
+ | `finish(options?)` | Stops recovery and carries on parsing from the current state |
272
+ | `insertToken(token)` | Shifts a token the input did not have, e.g. a synthesized `;` |
273
+ | `isSafe()` | True when the current state can consume the next real token |
274
+
275
+ ```ts
276
+ const parser = parserGenerator.generate(lexer, {
277
+ reducer: (_, { input }) => new Node(input),
278
+ recover({ lexer, states, statesStack, insertToken, addError, isSafe, finish, crash }) {
279
+ const token = lexer.peekNextToken();
280
+
281
+ // skip the extra semicolons the author left behind
282
+ while (token.type === TokenType.SEMICOLON) {
283
+ lexer.getNextToken();
284
+ if (isSafe()) return finish();
285
+ }
286
+
287
+ // or insert the semicolon they forgot
288
+ if (states[statesStack.peek()].getTerminalAction("SEMICOLON") != null) {
289
+ const semicolon = new Token(TokenType.SEMICOLON, ";", new ColumnAndRow(token.line, token.column), {});
290
+ const inserted = insertToken(semicolon);
291
+ if (inserted != null) return crash(inserted.reason);
292
+ addError(`Expected SEMICOLON but received ${TokenType[token.type]}`);
293
+ if (isSafe()) return finish();
294
+ }
295
+
296
+ return crash("Invalid syntax");
297
+ },
298
+ });
299
+
300
+ const { result, errors } = parser.parse();
301
+ ```
302
+
303
+ A parser consumes its lexer, so build a new one per input with `generate(...)`, or call `parser.reset()`
304
+ if you want to reuse the parser with a lexer that still has tokens left.
305
+
306
+ ## Generating and shipping the parsing table
307
+
308
+ Generating states for a large grammar takes a while (a few seconds for a few hundred productions), so
309
+ generate the table once at build time and commit it:
310
+
311
+ ```
312
+ npx sparse --input=grammar.txt --output=table.txt
313
+ ```
314
+
315
+ | Flag | Meaning |
316
+ | ----------------- | ----------------------------------------------------------------- |
317
+ | `--input=<file>` | Grammar to read (required) |
318
+ | `--output=<file>` | Table to write; missing directories are created |
319
+ | `--lalr` | Generate LALR(1) states instead of LR(1) |
320
+ | `--check` | Write nothing; verify that `<output>` already matches the grammar |
321
+ | `--stdout` | Print the table instead of writing it |
322
+ | `--quiet` | Hide warnings about suspicious rules |
323
+ | `--help` | Usage |
324
+
325
+ By default, LR(1) states are generated. You can generate LALR(1) states instead by passing the `--lalr`
326
+ flag (or setting `mode: "lalr1"` in `Sparse.fromProductions`): the resulting table is never larger than
327
+ the LR(1) one. Using LALR(1) yields a ~35% performance boost over LR(1) for the same grammar (tested on
328
+ the LoLang example).
329
+
330
+ `--check` is useful in CI to prove that a committed table still matches its grammar:
331
+
332
+ ```
333
+ npx sparse --input=example/math/grammar.txt --output=example/math/table.txt --check
334
+ ```
335
+
336
+ Then load the table with `Sparse.fromGrammarFile`, as shown in [Getting started](#getting-started).
337
+
338
+ ## API reference
339
+
340
+ ### Reading and writing grammars
341
+
342
+ | Function | Returns |
343
+ | ---------------------------------------------------------- | -------------------------------------------------------------------------------------- |
344
+ | `buildProductions(source)` / `tryBuildProductions(source)` | The unrolled productions. Throws / returns a `Result` failure with the offending token |
345
+ | `loadGrammar(path)` | Same, but reads the file, turning missing files into failures too |
346
+ | `buildStates(table)` / `tryBuildStates(table)` | `TableState[]` from a table file, validating action syntax and state numbers |
347
+ | `loadTable(path)` | Same, but reads the file |
348
+ | `generateStates(productions, options?)` | `GeneratorResult`; `options` is `{ mode, onWarning, onProgress }` |
349
+ | `validateProductions(productions, options?)` | Fails on a grammar that cannot generate a correct table |
350
+ | `validateTable(productions, states)` | Fails when a table does not belong to the productions |
351
+ | `validateTableStates(states, options?)` | Fails on a malformed table |
352
+
353
+ ### Building parsers
354
+
355
+ | Function | Returns |
356
+ | --------------------------------------------------------------------------------------------------- | ----------------------------------------------------------- |
357
+ | `Sparse.fromGrammarFile({ grammarPath, tablePath, toStringifiedTokenType, onWarning?, validate? })` | Loads a grammar and its prebuilt table, cross-checking them |
358
+ | `Sparse.fromProductions({ productions, toStringifiedTokenType, mode?, quiet? })` | Generates states at startup |
359
+ | `Sparse.tryFromProductions(...)` | Same, as a `Result` |
360
+ | `new Sparse({ productions, states, toStringifiedTokenType, validate?, source? })` | Use when you already hold the states |
361
+ | `generator.generate(lexer, { reducer, recover? })` | A parser |
362
+ | `parser.parse()` | `{ result, errors }` |
363
+ | `parser.reset()` | Clears the stacks and errors |
364
+
365
+ ### Helpers
366
+
367
+ | Export | Purpose |
368
+ | --------------------------------------------------------------------------- | ---------------------------------------------------------------- |
369
+ | `enumToString(TokenType)` | The `toStringifiedTokenType` every example used to write by hand |
370
+ | `defineReducers(reducers)` | Turns a map of named reducers into a reducer |
371
+ | `assertReducersCoverGrammar(productions, reducers)` | Fails up front when a named production has no reducer |
372
+ | `missingReducerNames(productions, reducers)` / `missingReducerMessage(...)` | Same, as data |
373
+ | `namedProductions(productions)` | Every production the grammar names |
374
+ | `TableState.fromJSObject(json)` / `GeneratorResult.toJSObject()` | Round trip a table through JSON |
375
+ | `hydrateProduction(json)` / `dehydateProduction(production)` | Round trip productions through JSON |
376
+ | `buildErrorWindow(source, token)` | Renders the window around a token |
377
+ | `Stack` | `push`, `pop`, `peek`, `peekAt`, `size`, `isEmpty`, `toArray` |
378
+
379
+ ## Known limitations
380
+
381
+ - `?`, `*` and `+` only apply to a parenthesized group, and a group that can match nothing
382
+ (`<A: a>: ([X])?;`) is rejected because Sparse has no support for empty productions. Write two
383
+ productions instead.
384
+ - Sparse does not detect grammar conflicts (shift/reduce, reduce/reduce). An ambiguous grammar produces a
385
+ table with one arbitrary resolution, so if your parser behaves strangely, check your grammar for ambiguity.
386
+ - A state with no actions cannot be written in the table file format; `toTable()` throws for such grammars.
387
+ Use `Sparse.fromProductions` to keep the states in memory instead.
388
+
129
389
  ## AI Disclaimer
130
390
 
131
391
  This project was originally written without the use of AI tools, the core LR(1) table generator was written by hand as per the algorithms described in the Dragon Book. Every release prior to v0.1 contained no AI generated code.
132
392
 
133
- OpenCode and Qwen 3.8 27B were to implement performance improvements on the original LR(1) table generator and to implement LALR(1) support.
393
+ OpenCode and Qwen 3.8 27B were to implement performance improvements on the original LR(1) table generator and to implement LALR(1) support.