metchurial 0.1.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (46) hide show
  1. metchurial-0.1.0/.gitignore +8 -0
  2. metchurial-0.1.0/LICENSE +21 -0
  3. metchurial-0.1.0/PKG-INFO +343 -0
  4. metchurial-0.1.0/README.md +325 -0
  5. metchurial-0.1.0/docs/PROVENANCE.md +151 -0
  6. metchurial-0.1.0/pyproject.toml +52 -0
  7. metchurial-0.1.0/src/metchurial/__init__.py +56 -0
  8. metchurial-0.1.0/src/metchurial/_generated/Db2Lexer.interp +2770 -0
  9. metchurial-0.1.0/src/metchurial/_generated/Db2Lexer.py +5021 -0
  10. metchurial-0.1.0/src/metchurial/_generated/Db2Lexer.tokens +1821 -0
  11. metchurial-0.1.0/src/metchurial/_generated/Db2Parser.interp +2871 -0
  12. metchurial-0.1.0/src/metchurial/_generated/Db2Parser.py +103350 -0
  13. metchurial-0.1.0/src/metchurial/_generated/Db2Parser.tokens +1821 -0
  14. metchurial-0.1.0/src/metchurial/_generated/Db2ParserVisitor.py +5153 -0
  15. metchurial-0.1.0/src/metchurial/_generated/__init__.py +0 -0
  16. metchurial-0.1.0/src/metchurial/cli.py +280 -0
  17. metchurial-0.1.0/src/metchurial/detect/__init__.py +5 -0
  18. metchurial-0.1.0/src/metchurial/detect/bad_file_check.py +78 -0
  19. metchurial-0.1.0/src/metchurial/detect/comment_rescan.py +121 -0
  20. metchurial-0.1.0/src/metchurial/detect/extractor_visitor.py +206 -0
  21. metchurial-0.1.0/src/metchurial/detect/supplementary_checks.py +203 -0
  22. metchurial-0.1.0/src/metchurial/engine.py +419 -0
  23. metchurial-0.1.0/src/metchurial/io_utils.py +143 -0
  24. metchurial-0.1.0/src/metchurial/mask.py +171 -0
  25. metchurial-0.1.0/src/metchurial/models/__init__.py +24 -0
  26. metchurial-0.1.0/src/metchurial/models/findings.py +30 -0
  27. metchurial-0.1.0/src/metchurial/models/identity.py +35 -0
  28. metchurial-0.1.0/src/metchurial/models/options.py +64 -0
  29. metchurial-0.1.0/src/metchurial/models/references.py +41 -0
  30. metchurial-0.1.0/src/metchurial/models/relations.py +44 -0
  31. metchurial-0.1.0/src/metchurial/models/results.py +48 -0
  32. metchurial-0.1.0/src/metchurial/models/tables.py +77 -0
  33. metchurial-0.1.0/src/metchurial/parsing/__init__.py +4 -0
  34. metchurial-0.1.0/src/metchurial/parsing/predicates.py +75 -0
  35. metchurial-0.1.0/src/metchurial/parsing/statement_driver.py +295 -0
  36. metchurial-0.1.0/src/metchurial/parsing/token_walk.py +35 -0
  37. metchurial-0.1.0/src/metchurial/references/__init__.py +5 -0
  38. metchurial-0.1.0/src/metchurial/references/function_visitor.py +113 -0
  39. metchurial-0.1.0/src/metchurial/references/query_identity.py +300 -0
  40. metchurial-0.1.0/src/metchurial/references/reference_visitor.py +57 -0
  41. metchurial-0.1.0/src/metchurial/references/relations.py +134 -0
  42. metchurial-0.1.0/src/metchurial/references/table_scan.py +620 -0
  43. metchurial-0.1.0/src/metchurial/report.py +393 -0
  44. metchurial-0.1.0/src/metchurial/split/__init__.py +2 -0
  45. metchurial-0.1.0/src/metchurial/split/select_blocks.py +124 -0
  46. metchurial-0.1.0/src/metchurial/tsv.py +39 -0
@@ -0,0 +1,8 @@
1
+ .DS_Store
2
+ __pycache__/
3
+ *.pyc
4
+ .venv/
5
+ .idea/
6
+ .ruff_cache/
7
+ gen/
8
+ pypi-dist/
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 CynicDog
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
@@ -0,0 +1,343 @@
1
+ Metadata-Version: 2.4
2
+ Name: metchurial
3
+ Version: 0.1.0
4
+ Summary: ANTLR-backed engine that turns DB2 SQL source into structured, queryable metadata -- with hardcoded-sensitive-value detection as one built-in analysis
5
+ Project-URL: Repository, https://github.com/CynicDog/metchurial
6
+ Project-URL: Issues, https://github.com/CynicDog/metchurial/issues
7
+ License-Expression: MIT
8
+ License-File: LICENSE
9
+ Classifier: Development Status :: 4 - Beta
10
+ Classifier: Intended Audience :: Developers
11
+ Classifier: Programming Language :: Python :: 3
12
+ Classifier: Programming Language :: Python :: 3 :: Only
13
+ Classifier: Topic :: Database
14
+ Classifier: Topic :: Software Development :: Quality Assurance
15
+ Requires-Python: >=3.9
16
+ Requires-Dist: antlr4-python3-runtime==4.13.2
17
+ Description-Content-Type: text/markdown
18
+
19
+ # metchurial
20
+
21
+ An ANTLR-backed static analysis engine for DB2 SQL: it parses SQL source
22
+ into a real syntax tree — not regex — using
23
+ [`antlr/grammars-v4`](https://github.com/antlr/grammars-v4)'s `sql/db2`
24
+ grammar, a purpose-built IBM Db2 LUW SQL grammar, and turns that tree into
25
+ structured, queryable metadata: every table, column, function, and
26
+ predicate reference a file makes, and the JOIN relationships between
27
+ tables, aggregated across an entire codebase. Legacy enterprise SQL is
28
+ itself a dataset worth analyzing systematically, not just code to review
29
+ file by file — that's what this toolkit is for. Hardcoded-sensitive-value
30
+ detection, SELECT-block splitting, and literal masking are all built on
31
+ that same parse tree, as specific analyses layered on top of it.
32
+
33
+ ## Capabilities
34
+
35
+ | Capability | Flag | Details |
36
+ |---|---|---|
37
+ | **Metadata extraction** — every table/column/function/predicate reference, JOIN relationships aggregated across the whole scan | `--extract-metadata` | [What it extracts](#what-it-extracts) |
38
+ | **Sensitive-value detection** — sensitive-column comparisons and known-name literals | *(default, always on)* | [What it detects](#what-it-detects) |
39
+ | **File splitting** — one file per standalone SELECT block | `--split-selects` | [Output artifacts](#output-artifacts) |
40
+ | **Literal masking** — rewrite flagged literals to fixed placeholders in place | `--mask-literals` | [Output artifacts](#output-artifacts) |
41
+
42
+ ## Quick start
43
+
44
+ You only need **one file**: `dist/metchurial.py`. It's a
45
+ self-contained bundle — the generated Db2 SQL parser and the ANTLR Python
46
+ runtime are both inlined into it, so it runs with plain `python` and no
47
+ `pip install`. This is verified in CI by running it in a virtualenv with
48
+ zero third-party packages installed (not even `antlr4-python3-runtime`)
49
+ and diffing the output against a normal run — see `tests/test_bundle.py`.
50
+
51
+ ```bat
52
+ python metchurial.py C:\sql\root
53
+ python metchurial.py C:\sql\root --sensitive-columns ACCT_ID CTRT_NO HLDR_NM
54
+ python metchurial.py C:\sql\root --extract-metadata --split-selects
55
+ python metchurial.py C:\sql\root --mask-literals
56
+ python metchurial.py C:\sql\root --workers 8 --verbose
57
+ ```
58
+
59
+ Every run writes its artifacts (`summary.md`, `findings.tsv`, ...) into the
60
+ **current working directory**, not the scan root — see
61
+ [Output artifacts](#output-artifacts) below. Exit code is `1` if anything
62
+ was found (FINDING), `0` if clean — convenient for wiring into a CI
63
+ step or a pre-commit check.
64
+
65
+ **Before carrying `dist/metchurial.py` into a restricted
66
+ environment**, read [How the bundle works](#how-the-bundle-works) below and
67
+ get `docs/PROVENANCE.md` reviewed by whoever handles third-party-code
68
+ intake there.
69
+
70
+ ### Running from source with uv
71
+
72
+ If you have [uv](https://docs.astral.sh/uv/) and don't need the
73
+ zero-dependency single-file artifact, you can run straight from source
74
+ instead of `dist/metchurial.py`:
75
+
76
+ ```bash
77
+ uv sync
78
+ uv run metchurial /sql/root
79
+ uv run metchurial /sql/root --extract-metadata --split-selects
80
+ ```
81
+
82
+ `uv sync` pulls in just the one runtime dependency this needs
83
+ (`antlr4-python3-runtime`, exact-pinned to match the version `src/metchurial/_generated/`'s
84
+ parser was built with) into a project-local `.venv` — no Java, no ANTLR
85
+ tooling required for that (those are dev-only, for regenerating
86
+ `src/metchurial/_generated/` or rebuilding `dist/metchurial.py` itself — see
87
+ [Dev workflow](#dev-workflow)).
88
+
89
+ ### Installing from PyPI
90
+
91
+ ```bash
92
+ pip install metchurial # or: uv add metchurial
93
+ ```
94
+
95
+ This installs the `metchurial` CLI and the library API below. The
96
+ single-file `dist/metchurial.py` remains the distribution channel for
97
+ restricted environments where even `pip install` isn't an option.
98
+
99
+ ### Using as a library
100
+
101
+ The CLI is one consumer of a plain Python API — install the package and
102
+ drive scans from your own code:
103
+
104
+ ```python
105
+ import metchurial
106
+
107
+ # One call: scan a tree with every metadata analysis on.
108
+ result = metchurial.scan("/sql/root", metchurial.ScanOptions.metadata())
109
+
110
+ for f in result.findings: # sensitive-value findings
111
+ print(f.file, f.line, f.column_name, f.value)
112
+ for row in result.identity_rows: # per-statement core_ids
113
+ print(row.core_id, row.file, row.line)
114
+
115
+ # Or per-file, with only what you need switched on:
116
+ one = metchurial.scan_file(
117
+ "query.sql",
118
+ metchurial.ScanOptions(sensitive_columns=("ACCT_ID", "HLDR_NM"),
119
+ extract_relations=True))
120
+ ```
121
+
122
+ `scan()`/`scan_file()` return typed result objects
123
+ (`TreeScanResult`/`FileScanResult` of `Finding`/`TableUse`/`RelationEdge`/
124
+ `IdentityRow`/... rows) and never print or write files on their own —
125
+ report artifacts (`summary.md`, `findings.tsv`, `refs_*.tsv`) are the
126
+ CLI's job. The one exception is `ScanOptions(split_selects=True)`, which
127
+ writes `-NN` split files next to each multi-SELECT source file, same as
128
+ the `--split-selects` flag. Everything a scan can be told is a field on
129
+ `ScanOptions` (a frozen dataclass); `ScanOptions.metadata()` is shorthand
130
+ for switching every `extract_*` analysis on, mirroring
131
+ `--extract-metadata`.
132
+
133
+ ## CLI reference
134
+
135
+ | Flag | Default | Description |
136
+ |---|---|---|
137
+ | `root` (positional) | — | Directory to scan recursively |
138
+ | `--sensitive-columns` | `ACCT_ID CTRT_NO ACCT_NM ACCT_NAME` | Column names sensitive-column comparison detection treats as sensitive; fully replaces the default list, doesn't add to it |
139
+ | `--extensions` | `sql txt` | File extensions to scan, without the dot |
140
+ | `--extract-metadata` | off | Also emit `refs_tables.tsv`/`refs_columns.tsv`/`refs_functions.tsv`/`refs_relations.tsv`/`refs_query_identity.tsv` (schema/table/column refs, JOIN relationships, function/predicate usage, per-statement structural identity) and matching summary.md sections — see [Output artifacts](#output-artifacts) |
141
+ | `--query-similarity` | off | Also emit `refs_query_similarity.tsv`: pairwise Jaccard similarity between statements that don't share a `core_id`. Opt-in because the pass is O(n²) in the number of *distinct* core_ids — fine for thousands of distinct queries, slow for tens of thousands. Requires `--extract-metadata` |
142
+ | `--split-selects` | off | For a file with 2+ standalone SELECT blocks, write one `<stem>-NN<ext>` file per block alongside the original (files with a single block are left as-is) |
143
+ | `--mask-literals` | off | Rewrite in place every flagged literal's content to a fixed placeholder (`'****'`/`"****"` for quoted, `0000` for unquoted numeric), everything else byte-for-byte identical — back up files first, this overwrites them |
144
+ | `--workers N` | `1` | Scan across N worker processes instead of one |
145
+ | `--max-chunk-iterations N` | `200000` | Safety-valve cap on the resync driver's loop iterations per statement chunk |
146
+ | `--verbose` | off | Print a `[i/N]` progress line to stderr as each file is scanned |
147
+
148
+ `--workers N` scans files across N worker processes (`concurrent.futures.
149
+ ProcessPoolExecutor`, stdlib only). Parsing is CPU-bound pure Python, so
150
+ this is real multi-core parallelism, not threads, which the GIL would keep
151
+ from helping here. Each file is scanned independently with no shared
152
+ state, so results are unaffected other than which order they're merged in
153
+ (the reports already group/sort by file and line regardless). Leave a
154
+ couple of cores free for the OS/other work rather than setting `--workers`
155
+ to your full core count.
156
+
157
+ ## What it extracts
158
+
159
+ `--extract-metadata` walks the same parse tree detection uses, but
160
+ unconditionally — every reference in the file, not just ones compared to a
161
+ literal — and resolves each one back to the table/schema it actually
162
+ belongs to:
163
+
164
+ - **Table & schema references** — every `schema.table` a file's SQL
165
+ touches. Each query block gets its own alias map, so a bare `t1` or
166
+ `t1.col` resolves back to the schema-qualified table it actually refers
167
+ to, not just the identifier as written on that line.
168
+ - **Column references** — every `schema.table.column` reference in the
169
+ file, not only ones inside a comparison; a correlated subquery's own
170
+ `t.col` is resolved within its own scope, not leaked into its parent's.
171
+ - **Function & predicate usage** — every function call (`SUBSTR`,
172
+ `COALESCE`, `SUM`, ...) and predicate operator (`=`, `IN`, `BETWEEN`,
173
+ `LIKE`, `IS NULL`, ...) actually used, with the exact source text of its
174
+ arguments/operands.
175
+ - **JOIN relationships** — every table-to-table JOIN edge, aggregated
176
+ across the *entire scan* (one graph, not one per file) — how tables in a
177
+ legacy schema actually connect in practice, not what an ER diagram
178
+ claims they should.
179
+
180
+ Each of these lands in its own `refs_*.tsv` — see
181
+ [Output artifacts](#output-artifacts) — built to be loaded straight into a
182
+ spreadsheet or a graph tool.
183
+
184
+ ## What it detects
185
+
186
+ Sensitive-value detection is one specific analysis built on the same parse
187
+ tree metadata extraction uses — always on, independent of
188
+ `--extract-metadata`. It's two independent mechanisms, each producing its
189
+ own finding:
190
+
191
+ - **Sensitive-Column Comparison Detection — FINDING**: a sensitive column
192
+ (see `--sensitive-columns`) compared to a hardcoded literal — `=`, `<>`,
193
+ `!=`, `<=`, `>=`, `<`, `>`, `(NOT) IN (...)`, `(NOT) LIKE`,
194
+ `BETWEEN ... AND ...`, or a bare `(` before a literal (a DB2 quirk) — in
195
+ either direction and regardless of spacing or line breaks.
196
+ - **Known-Name Matching — FINDING**: any quoted, name-shaped literal (2-4
197
+ Hangul syllables) whose text is listed in `known_names.txt`, regardless
198
+ of which column it's compared to. There's no surname heuristic — a
199
+ literal only becomes a finding once a human has confirmed it's a real
200
+ name. Every other name-shaped literal is a triage *candidate*: it shows up
201
+ in `strings.txt` each run until it's copied into either `known_names.txt`
202
+ (flags it as a finding from then on) or `stopwords.txt` (excludes it from
203
+ `strings.txt` from then on). Both files are one word per line, `#`
204
+ comments allowed, auto-created empty on first run — see
205
+ [Output artifacts](#output-artifacts).
206
+
207
+ Findings inside `--`/`/* */` comments are still reported (commented-out
208
+ code can leak real data) and tagged `in_comment=Y` in `findings.tsv`.
209
+ `/* */` comments may nest, and a finding inside a nested comment is still
210
+ found.
211
+
212
+ ### Public placeholder names
213
+
214
+ The column and table names used throughout this repo — `ACCT_ID`,
215
+ `CTRT_NO`, `ACCT_NM`, `ACCT_NAME`, `HLDR_NM`, `TBSAMPLE001`, `STAT_CD` —
216
+ are placeholders, not real production schema names from any actual DB2
217
+ environment, swapped in consistently before this repo was made public.
218
+ They appear in `DEFAULT_SENSITIVE_COLUMNS` (`src/metchurial/models/options.py`, `--sensitive-columns`'s
219
+ built-in default), every fixture under `tests/fixtures/`, and this
220
+ README's own examples. This doesn't affect behavior — `--sensitive-columns`
221
+ always fully replaces the default list, so a real deployment should pass
222
+ its own actual column names explicitly on every run rather than relying on
223
+ the shipped defaults meaning anything for your schema.
224
+
225
+ ## Output artifacts
226
+
227
+ Every scan writes the same fixed set of files into the current working
228
+ directory (not the scan root, and not configurable — one predictable set
229
+ of names regardless of invocation). `summary.md` is an index into the
230
+ others: bounded counts and top-N tables with pointers to the full detail,
231
+ not a duplicate of it.
232
+
233
+ | File | Written when | Contents |
234
+ |---|---|---|
235
+ | `summary.md` | always | Run info, Sensitive Findings (with per-file detail), String Occurrences, Bad Files, Stopwords, Known Names, and — if enabled — Table & Column References, Functions, Relations, Select Blocks |
236
+ | `findings.tsv` | always | Every finding, one row per literal, for filtering/sorting in Excel |
237
+ | `strings.txt` | always | Unique name-shaped literals not yet classified into `known_names.txt`/`stopwords.txt`, with occurrence counts, in a format directly copy-pasteable into either |
238
+ | `stopwords.txt` | always (auto-created empty with a format header if missing) | Name-shaped literals reviewed and confirmed *not* sensitive — excluded from `strings.txt` from then on; edit in place |
239
+ | `known_names.txt` | always (auto-created empty with a format header if missing) | Name-shaped literals reviewed and confirmed sensitive — every matching literal becomes a known-name finding from then on; edit in place |
240
+ | `bad_files.txt` | always | Persistent skip-list of files too malformed to parse — see [Bad files](#bad-files) |
241
+ | `refs_tables.tsv` | `--extract-metadata` | Every `schema.table` reference found, with file/line |
242
+ | `refs_columns.tsv` | `--extract-metadata` | Every `schema.table.column` reference found, with file/line |
243
+ | `refs_functions.tsv` | `--extract-metadata` | Every function call and predicate operator found, with operands/file/line |
244
+ | `refs_relations.tsv` | `--extract-metadata` | Table-to-table JOIN usage aggregated across the whole scan (one file, not per-directory) |
245
+ | `refs_query_identity.tsv` | `--extract-metadata` | One `core_id` per statement — structurally identical statements share one id regardless of aliasing/projection/formatting differences |
246
+ | `refs_query_similarity.tsv` | `--query-similarity` | Pairwise Jaccard similarity between distinct `core_id`s that don't match exactly |
247
+
248
+ ## Bad files
249
+
250
+ Some real-world SQL files aren't really valid SQL — internal section
251
+ dividers (`========`, `<<목표KPI>>`), bare prose headers, Korean-language
252
+ comments used as informal headings, or files with the actual SQL truncated
253
+ partway through. These can make the parser resync loop grind for a very
254
+ long time on a single file, or crash it outright.
255
+
256
+ Two independent safety nets guard against this:
257
+
258
+ - A cheap **pre-check** on the token stream (lexer-error ratio and
259
+ long runs of repeated punctuation) flags a file as bad before any real
260
+ parsing is attempted.
261
+ - A broad **try/except** around the actual scan of each file catches any
262
+ unexpected crash and treats it the same way.
263
+
264
+ Either path records the file's path and a short reason in `bad_files.txt`,
265
+ and the file is skipped entirely — not even attempted — on every later
266
+ run. Workflow:
267
+
268
+ 1. Run the scan; anything unfixably weird lands in `bad_files.txt` and is
269
+ skipped from then on.
270
+ 2. Fix the file's SQL content by hand (or decide it's fine to leave out).
271
+ 3. Delete that file's line from `bad_files.txt`.
272
+ 4. Re-run — the file is attempted again on the next scan.
273
+
274
+ `bad_files.txt` is a local, per-environment artifact (it's `.gitignore`d)
275
+ rather than something meant to be committed and shared.
276
+
277
+ ## Known limitations
278
+
279
+ Known gaps in the vendored grammar, `--extract-metadata` extraction, and
280
+ the file-encoding auto-detection are tracked as GitHub issues, not
281
+ duplicated here — refer to
282
+ [Issues](https://github.com/CynicDog/metchurial/issues). Each is backed
283
+ by a runnable test (`tests/test_grammar_smoke.py`,
284
+ `tests/test_db2_grammar_specific_cases.py`, and friends), not just prose.
285
+
286
+ ## How the bundle works
287
+
288
+ `dist/metchurial.py` is built with
289
+ [stickytape](https://github.com/mwilliamson/stickytape), which inlines
290
+ every module's source as embedded strings. **At every run**, it writes
291
+ that source out to a fresh OS temp directory, imports from there, and
292
+ deletes it on exit. Practically:
293
+
294
+ - It needs write access to the OS temp directory (normally fine, but worth
295
+ confirming in advance).
296
+ - Some corporate EDR/antivirus tools are wary of "a script writes many
297
+ `.py` files to temp and imports them" as a pattern. If your environment
298
+ has an infra security review step, flag this mechanism to them up front
299
+ — `docs/PROVENANCE.md` documents exactly what's bundled and why.
300
+
301
+ The file is large (~4.8MB, mostly the Db2 parser's serialized ATN tables)
302
+ -- one file, but not a small one, and not realistically human-auditable
303
+ top to bottom.
304
+
305
+ ## Dev workflow
306
+
307
+ Everything below needs the dev dependency group (`antlr4-tools`,
308
+ `stickytape`, `ruff` — see `pyproject.toml`) plus Java for grammar
309
+ regeneration specifically (`uv`/`pip` can't install that) — none of this
310
+ runs in the restricted target environment; only `dist/metchurial.py` does.
311
+
312
+ ```bash
313
+ uv sync
314
+
315
+ # regenerate src/metchurial/_generated/ from vendor/grammars-v4/*.g4 (only needed after
316
+ # touching the grammar itself)
317
+ uv run bash build/generate_parser.sh
318
+
319
+ # run everything: grammar smoke tests, scan_file()-level tests, end-to-end
320
+ # edge-case regressions, and the bundle self-containment check
321
+ uv run python -m unittest discover -s tests -p "test_*.py"
322
+
323
+ # lint
324
+ uv run ruff check src
325
+
326
+ # rebuild the deployable single-file artifact
327
+ uv run python build/bundle.py
328
+
329
+ # build the PyPI sdist+wheel (kept out of dist/, which holds the bundle)
330
+ uv build --out-dir pypi-dist
331
+ ```
332
+
333
+ Publishing to PyPI is automated: publishing a GitHub release triggers
334
+ `.github/workflows/publish.yml`, which builds and uploads via PyPI
335
+ trusted publishing — no API tokens involved.
336
+
337
+ ## Licensing
338
+
339
+ This project vendors two pieces of third-party code into the deployable
340
+ artifact: an IBM Db2 SQL grammar (MIT, from `antlr/grammars-v4`'s
341
+ `sql/db2`) and `antlr4-python3-runtime` (BSD-3-Clause). See
342
+ `docs/PROVENANCE.md` for exact versions, license texts, and what (if
343
+ anything) was modified.