@cirru/parser.ts 0.0.8 → 0.0.9

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Binary file
package/.yarnrc.yml ADDED
@@ -0,0 +1 @@
1
+ nodeLinker: node-modules
package/Agents.md ADDED
@@ -0,0 +1,198 @@
1
+ # Agents.md — Development, Testing & Optimization Notes
2
+
3
+ Quick-reference for contributors and AI agents working on this codebase.
4
+
5
+ ---
6
+
7
+ ## Project Overview
8
+
9
+ `parser.ts` is a TypeScript implementation of the [Cirru](http://cirru.org/) indentation-sensitive parser.
10
+ It converts Cirru source text into a nested `ICirruNode` (string | ICirruNode[]) tree.
11
+
12
+ Main exports: `parse(code)` and `parseOneLiner(code)` from `src/index.ts`.
13
+
14
+ ---
15
+
16
+ ## Quick Commands
17
+
18
+ ```bash
19
+ yarn test # run all 22 Jest tests
20
+ yarn bench # benchmark on bundled test fixtures (×20, ~16 KB)
21
+ BENCH_FILE=/absolute/path/to/large.cirru yarn bench # benchmark on a real file
22
+ yarn compile # emit lib/ (tsc -p tsconfig-compile.json)
23
+ yarn release # production bundle via vite (rolldown)
24
+ ```
25
+
26
+ ---
27
+
28
+ ## Package Manager
29
+
30
+ Yarn Berry 4.x with `nodeLinker: node-modules` (`.yarnrc.yml`).
31
+
32
+ ```bash
33
+ corepack enable # one-time setup per machine
34
+ yarn install --immutable # CI; exact lockfile
35
+ ```
36
+
37
+ CI workflows (`.github/workflows/`) enable Corepack before running Yarn commands.
38
+ Do **not** use `actions/setup-node` with `cache: yarn` here, because its cache probe runs before `corepack enable` and can fail by invoking global Yarn 1 against the project's `packageManager: yarn@4.13.0`.
39
+
40
+ ---
41
+
42
+ ## Testing
43
+
44
+ - Framework: Jest + `ts-jest`
45
+ - Test file: `src/parser.test.ts`
46
+ - Fixtures: `test/cirru/*.cirru` (source) + `test/ast/*.json` (expected output)
47
+ - Every optimization must keep all **27 tests passing**.
48
+
49
+ Rules:
50
+
51
+ - Run `yarn test` before and after any structural change.
52
+ - `folded-beginning.cirru` is the trickiest edge case: the file can begin with a blank line and then indented content, so the very first indentation flush must preserve top-level structure instead of introducing an extra sibling boundary.
53
+ - Any optimization that touches `$` or `,` semantics must be treated as high-risk and validated against the full fixture suite before considering benchmark wins.
54
+ - There are now explicit tree-level equivalence tests for the combined `$`/`,` pass; keep them green before trusting any transform rewrite.
55
+
56
+ ---
57
+
58
+ ## Benchmark Setup
59
+
60
+ `src/bench.ts` compiles with `tsconfig.bench.json` (target es2022, skipLibCheck, includes node types),
61
+ then runs with plain `node`.
62
+
63
+ ```bash
64
+ # Benchmark internals
65
+ iterations = 500 (with 50 warmup rounds)
66
+ BENCH_FILE env var → absolute path, never committed (large real-world file)
67
+ fallback → test/cirru/*.cirru joined ×20 (~16 KB)
68
+ ```
69
+
70
+ ```bash
71
+ # Typical real-world numbers (2026-03, Apple Silicon)
72
+ # Small fixtures ×20 (16 KB): ~5,000 ops/sec (was ~2,250 at original baseline)
73
+ # calcit-core.cirru (244 KB): ~594 ops/sec (was ~323 at original baseline)
74
+ ```
75
+
76
+ ---
77
+
78
+ ## Architecture
79
+
80
+ ```
81
+ parse(code)
82
+ └─ lexAndBuild(code) # single-pass: lex + indent + tree-build + dollar/comma
83
+ stack-based tree builder, no intermediate token array
84
+ → ICirruNode[]
85
+ ```
86
+
87
+ ### ELexState machine (in `lexAndBuild`)
88
+
89
+ | State | Description |
90
+ | -------- | ------------------------------------- |
91
+ | `indent` | counting leading spaces at line start |
92
+ | `space` | between tokens |
93
+ | `token` | inside a bare word/symbol |
94
+ | `string` | inside a `"..."` literal |
95
+ | `escape` | after `\` inside a string |
96
+
97
+ ---
98
+
99
+ ## Optimization History & Techniques
100
+
101
+ ### Round 1 — in `lexAndResolve` (commit f9781a8)
102
+
103
+ - **Eliminate array destructuring** in hot loop: `[acc, state, buffer] = [acc, ...]` → direct assignments.
104
+ - **No template literals** in hot path: `\`${buffer}${c}\``→`buffer + c`.
105
+ - **Inline `repeat()`/`pushToList()`** → bare `for` loops.
106
+ - **Cache `code.length`** into `len` before the loop.
107
+ - Result: +34% on small input.
108
+
109
+ ### Round 2 — merged passes (commit c441701)
110
+
111
+ - **Single-pass lex + indent resolution**: eliminated the intermediate `LexList` array that the old `lex()` → `resolveIndentations()` two-step created.
112
+ - **Integer indent counter**: `indentCount++` instead of string `" ".repeat(n)` comparison.
113
+ - **`hasDollar` / `hasComma` flags**: skip expensive tree-transform passes entirely when the source contains neither `$` nor `,` tokens.
114
+
115
+ ### Round 3 — slice-based token extraction (commit c441701)
116
+
117
+ - **`code.slice(tokenStart, pointer - 1)`** at token boundaries instead of `buffer + c` per character. Eliminates O(token_length) string allocations per token.
118
+ - **`stringStart` + `stringHasEscape` flag**: string literals use `code.slice(stringStart, pointer - 1)` when no backslash is encountered; only fall back to char-by-char `buffer +` when an escape sequence appears.
119
+
120
+ ### Round 4 — no-slice tree helpers (commit f7e4fb9)
121
+
122
+ - **`dollarHelper(after, start)`**: pass array + start index instead of `after.slice(pointer + 1)` on every `$` token.
123
+ - **`commaHelper(after, start)`**: same — avoids `cursor.slice(1)` allocation per comma-head expression.
124
+ - Result: +12% on 244 KB real-world file (353 → 395 ops/sec).
125
+
126
+ ### Round 5 — eliminate `tokens[]` array (commit 80ced1b)
127
+
128
+ - **`lexAndBuild(code)`**: merged `lexAndResolve` + `buildExprs` + `graspeExprs` into a single stack-based pass.
129
+ - No `tokens[]` array, no `pointer` index, no `pullToken` closure — all eliminated.
130
+ - Stack model: `result[]` (top-level output), `stack[]` (in-progress arrays), `current` (innermost).
131
+ - `emitOpen()` / `emitClose()` / `emitToken()` are the only three operations; called directly from the character loop.
132
+ - Outer `open` is emitted lazily on first non-whitespace (`first` flag), so end-of-input needs no `acc.unshift()`.
133
+ - Result: +16% small (3,118 → 3,615 ops/sec), **+21% large (395 → 479 ops/sec)**.
134
+ - **Total from original baseline**: +61% (16 KB) / +48% (244 KB).
135
+
136
+ ### Round 6 — numeric dispatch in hot lexer loop
137
+
138
+ - **`code.charCodeAt(pointer++)`** replaced `code[pointer++]` in `lexAndBuild()`.
139
+ - Hot-path state dispatch now compares integer char codes for space/newline/quote/parens/backslash instead of 1-char strings.
140
+ - `code.slice(...)` is still used for token extraction, so token contents and edge-case behavior stay unchanged.
141
+ - Escape handling still reconstructs exact characters, but only on the rare escape path.
142
+ - Result: **3,615 → 4,013 ops/sec** on 16 KB fixtures, **479 → 496 ops/sec** on 244 KB real-world input.
143
+
144
+ ### Round 7 — combine `$` and `,` tree transforms
145
+
146
+ - Added **`resolveDollarComma(xs)`** in `src/tree.ts` as a single recursive pass intended to be equivalent to `resolveComma(resolveDollar(xs))`.
147
+ - Parser now uses the combined transform whenever either `hasDollar` or `hasComma` is present.
148
+ - Added focused equivalence tests for empty input, comma expansion, dollar nesting, unfolding-style mixed input, and nested-array-head behavior.
149
+ - This is a semantics-sensitive optimization, so the extra tests are part of the safety net, not just regression coverage.
150
+ - Result: **4,013 → 5,058 ops/sec** on 16 KB fixtures, **496 → 594 ops/sec** on 244 KB real-world input.
151
+
152
+ ---
153
+
154
+ ## V8 Profiling
155
+
156
+ ```bash
157
+ ./node_modules/.bin/tsc -p tsconfig.bench.json --outDir .bench
158
+ node --prof .bench/bench.js
159
+ node --prof-process isolate-*.log 2>/dev/null | head -100
160
+ rm -f isolate-*.log .bench/
161
+ ```
162
+
163
+ Latest profile findings (2026-03, after round 7, large real-world file):
164
+ | Function | JS ticks |
165
+ |------------------|-----------|
166
+ | `lexAndBuild` | 58.3% |
167
+ | `dollarCommaHelper` | 12.7% |
168
+
169
+ Interpretation:
170
+
171
+ - The combined tree pass paid off: the two separate helpers disappeared from the profile and total GC dropped again.
172
+ - `lexAndBuild()` is now even more clearly the dominant pure-JS hotspot.
173
+ - Future lexer work should still prefer local, semantics-preserving changes unless a larger rewrite can be defended with tests first.
174
+
175
+ ---
176
+
177
+ ## Known Bottlenecks / Future Work
178
+
179
+ 1. **`lexAndBuild()` dispatch cost** — now ~58% of JS ticks on the large-file profile. This is the clearest remaining hotspot.
180
+ 2. **Character extraction on rare paths** — escape handling still uses `code[pointer - 1]`; tiny wins may still exist, but they should stay local and preserve exact string semantics.
181
+ 3. **Array pre-sizing / pooling** — possible, but lower priority than lexer work and easier to get wrong for little gain.
182
+
183
+ ---
184
+
185
+ ## File Map
186
+
187
+ | File | Purpose |
188
+ | ----------------------- | ----------------------------------------------------------- |
189
+ | `src/index.ts` | Parser entry point; `lexAndBuild`, `parse`, `parseOneLiner` |
190
+ | `src/tree.ts` | Tree helpers: `resolveDollar`, `resolveComma`, utilities |
191
+ | `src/types.ts` | `ELexState`, `ELexControl`, `ICirruNode` |
192
+ | `src/parser.test.ts` | All 22 tests |
193
+ | `src/bench.ts` | Benchmark script |
194
+ | `tsconfig.bench.json` | Separate tsconfig for bench (es2022, node types) |
195
+ | `tsconfig-compile.json` | Library build config |
196
+ | `test/cirru/*.cirru` | Test input fixtures |
197
+ | `test/ast/*.json` | Expected AST outputs |
198
+ | `lib/` | Compiled library output (committed) |