metchurial 0.1.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- metchurial-0.1.0/.gitignore +8 -0
- metchurial-0.1.0/LICENSE +21 -0
- metchurial-0.1.0/PKG-INFO +343 -0
- metchurial-0.1.0/README.md +325 -0
- metchurial-0.1.0/docs/PROVENANCE.md +151 -0
- metchurial-0.1.0/pyproject.toml +52 -0
- metchurial-0.1.0/src/metchurial/__init__.py +56 -0
- metchurial-0.1.0/src/metchurial/_generated/Db2Lexer.interp +2770 -0
- metchurial-0.1.0/src/metchurial/_generated/Db2Lexer.py +5021 -0
- metchurial-0.1.0/src/metchurial/_generated/Db2Lexer.tokens +1821 -0
- metchurial-0.1.0/src/metchurial/_generated/Db2Parser.interp +2871 -0
- metchurial-0.1.0/src/metchurial/_generated/Db2Parser.py +103350 -0
- metchurial-0.1.0/src/metchurial/_generated/Db2Parser.tokens +1821 -0
- metchurial-0.1.0/src/metchurial/_generated/Db2ParserVisitor.py +5153 -0
- metchurial-0.1.0/src/metchurial/_generated/__init__.py +0 -0
- metchurial-0.1.0/src/metchurial/cli.py +280 -0
- metchurial-0.1.0/src/metchurial/detect/__init__.py +5 -0
- metchurial-0.1.0/src/metchurial/detect/bad_file_check.py +78 -0
- metchurial-0.1.0/src/metchurial/detect/comment_rescan.py +121 -0
- metchurial-0.1.0/src/metchurial/detect/extractor_visitor.py +206 -0
- metchurial-0.1.0/src/metchurial/detect/supplementary_checks.py +203 -0
- metchurial-0.1.0/src/metchurial/engine.py +419 -0
- metchurial-0.1.0/src/metchurial/io_utils.py +143 -0
- metchurial-0.1.0/src/metchurial/mask.py +171 -0
- metchurial-0.1.0/src/metchurial/models/__init__.py +24 -0
- metchurial-0.1.0/src/metchurial/models/findings.py +30 -0
- metchurial-0.1.0/src/metchurial/models/identity.py +35 -0
- metchurial-0.1.0/src/metchurial/models/options.py +64 -0
- metchurial-0.1.0/src/metchurial/models/references.py +41 -0
- metchurial-0.1.0/src/metchurial/models/relations.py +44 -0
- metchurial-0.1.0/src/metchurial/models/results.py +48 -0
- metchurial-0.1.0/src/metchurial/models/tables.py +77 -0
- metchurial-0.1.0/src/metchurial/parsing/__init__.py +4 -0
- metchurial-0.1.0/src/metchurial/parsing/predicates.py +75 -0
- metchurial-0.1.0/src/metchurial/parsing/statement_driver.py +295 -0
- metchurial-0.1.0/src/metchurial/parsing/token_walk.py +35 -0
- metchurial-0.1.0/src/metchurial/references/__init__.py +5 -0
- metchurial-0.1.0/src/metchurial/references/function_visitor.py +113 -0
- metchurial-0.1.0/src/metchurial/references/query_identity.py +300 -0
- metchurial-0.1.0/src/metchurial/references/reference_visitor.py +57 -0
- metchurial-0.1.0/src/metchurial/references/relations.py +134 -0
- metchurial-0.1.0/src/metchurial/references/table_scan.py +620 -0
- metchurial-0.1.0/src/metchurial/report.py +393 -0
- metchurial-0.1.0/src/metchurial/split/__init__.py +2 -0
- metchurial-0.1.0/src/metchurial/split/select_blocks.py +124 -0
- metchurial-0.1.0/src/metchurial/tsv.py +39 -0
metchurial-0.1.0/LICENSE
ADDED
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
MIT License
|
|
2
|
+
|
|
3
|
+
Copyright (c) 2026 CynicDog
|
|
4
|
+
|
|
5
|
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
|
6
|
+
of this software and associated documentation files (the "Software"), to deal
|
|
7
|
+
in the Software without restriction, including without limitation the rights
|
|
8
|
+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
|
9
|
+
copies of the Software, and to permit persons to whom the Software is
|
|
10
|
+
furnished to do so, subject to the following conditions:
|
|
11
|
+
|
|
12
|
+
The above copyright notice and this permission notice shall be included in all
|
|
13
|
+
copies or substantial portions of the Software.
|
|
14
|
+
|
|
15
|
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
|
16
|
+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
|
17
|
+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
|
18
|
+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
|
19
|
+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
|
20
|
+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
|
21
|
+
SOFTWARE.
|
|
@@ -0,0 +1,343 @@
|
|
|
1
|
+
Metadata-Version: 2.4
|
|
2
|
+
Name: metchurial
|
|
3
|
+
Version: 0.1.0
|
|
4
|
+
Summary: ANTLR-backed engine that turns DB2 SQL source into structured, queryable metadata -- with hardcoded-sensitive-value detection as one built-in analysis
|
|
5
|
+
Project-URL: Repository, https://github.com/CynicDog/metchurial
|
|
6
|
+
Project-URL: Issues, https://github.com/CynicDog/metchurial/issues
|
|
7
|
+
License-Expression: MIT
|
|
8
|
+
License-File: LICENSE
|
|
9
|
+
Classifier: Development Status :: 4 - Beta
|
|
10
|
+
Classifier: Intended Audience :: Developers
|
|
11
|
+
Classifier: Programming Language :: Python :: 3
|
|
12
|
+
Classifier: Programming Language :: Python :: 3 :: Only
|
|
13
|
+
Classifier: Topic :: Database
|
|
14
|
+
Classifier: Topic :: Software Development :: Quality Assurance
|
|
15
|
+
Requires-Python: >=3.9
|
|
16
|
+
Requires-Dist: antlr4-python3-runtime==4.13.2
|
|
17
|
+
Description-Content-Type: text/markdown
|
|
18
|
+
|
|
19
|
+
# metchurial
|
|
20
|
+
|
|
21
|
+
An ANTLR-backed static analysis engine for DB2 SQL: it parses SQL source
|
|
22
|
+
into a real syntax tree — not regex — using
|
|
23
|
+
[`antlr/grammars-v4`](https://github.com/antlr/grammars-v4)'s `sql/db2`
|
|
24
|
+
grammar, a purpose-built IBM Db2 LUW SQL grammar, and turns that tree into
|
|
25
|
+
structured, queryable metadata: every table, column, function, and
|
|
26
|
+
predicate reference a file makes, and the JOIN relationships between
|
|
27
|
+
tables, aggregated across an entire codebase. Legacy enterprise SQL is
|
|
28
|
+
itself a dataset worth analyzing systematically, not just code to review
|
|
29
|
+
file by file — that's what this toolkit is for. Hardcoded-sensitive-value
|
|
30
|
+
detection, SELECT-block splitting, and literal masking are all built on
|
|
31
|
+
that same parse tree, as specific analyses layered on top of it.
|
|
32
|
+
|
|
33
|
+
## Capabilities
|
|
34
|
+
|
|
35
|
+
| Capability | Flag | Details |
|
|
36
|
+
|---|---|---|
|
|
37
|
+
| **Metadata extraction** — every table/column/function/predicate reference, JOIN relationships aggregated across the whole scan | `--extract-metadata` | [What it extracts](#what-it-extracts) |
|
|
38
|
+
| **Sensitive-value detection** — sensitive-column comparisons and known-name literals | *(default, always on)* | [What it detects](#what-it-detects) |
|
|
39
|
+
| **File splitting** — one file per standalone SELECT block | `--split-selects` | [Output artifacts](#output-artifacts) |
|
|
40
|
+
| **Literal masking** — rewrite flagged literals to fixed placeholders in place | `--mask-literals` | [Output artifacts](#output-artifacts) |
|
|
41
|
+
|
|
42
|
+
## Quick start
|
|
43
|
+
|
|
44
|
+
You only need **one file**: `dist/metchurial.py`. It's a
|
|
45
|
+
self-contained bundle — the generated Db2 SQL parser and the ANTLR Python
|
|
46
|
+
runtime are both inlined into it, so it runs with plain `python` and no
|
|
47
|
+
`pip install`. This is verified in CI by running it in a virtualenv with
|
|
48
|
+
zero third-party packages installed (not even `antlr4-python3-runtime`)
|
|
49
|
+
and diffing the output against a normal run — see `tests/test_bundle.py`.
|
|
50
|
+
|
|
51
|
+
```bat
|
|
52
|
+
python metchurial.py C:\sql\root
|
|
53
|
+
python metchurial.py C:\sql\root --sensitive-columns ACCT_ID CTRT_NO HLDR_NM
|
|
54
|
+
python metchurial.py C:\sql\root --extract-metadata --split-selects
|
|
55
|
+
python metchurial.py C:\sql\root --mask-literals
|
|
56
|
+
python metchurial.py C:\sql\root --workers 8 --verbose
|
|
57
|
+
```
|
|
58
|
+
|
|
59
|
+
Every run writes its artifacts (`summary.md`, `findings.tsv`, ...) into the
|
|
60
|
+
**current working directory**, not the scan root — see
|
|
61
|
+
[Output artifacts](#output-artifacts) below. Exit code is `1` if anything
|
|
62
|
+
was found (FINDING), `0` if clean — convenient for wiring into a CI
|
|
63
|
+
step or a pre-commit check.
|
|
64
|
+
|
|
65
|
+
**Before carrying `dist/metchurial.py` into a restricted
|
|
66
|
+
environment**, read [How the bundle works](#how-the-bundle-works) below and
|
|
67
|
+
get `docs/PROVENANCE.md` reviewed by whoever handles third-party-code
|
|
68
|
+
intake there.
|
|
69
|
+
|
|
70
|
+
### Running from source with uv
|
|
71
|
+
|
|
72
|
+
If you have [uv](https://docs.astral.sh/uv/) and don't need the
|
|
73
|
+
zero-dependency single-file artifact, you can run straight from source
|
|
74
|
+
instead of `dist/metchurial.py`:
|
|
75
|
+
|
|
76
|
+
```bash
|
|
77
|
+
uv sync
|
|
78
|
+
uv run metchurial /sql/root
|
|
79
|
+
uv run metchurial /sql/root --extract-metadata --split-selects
|
|
80
|
+
```
|
|
81
|
+
|
|
82
|
+
`uv sync` pulls in just the one runtime dependency this needs
|
|
83
|
+
(`antlr4-python3-runtime`, exact-pinned to match the version `src/metchurial/_generated/`'s
|
|
84
|
+
parser was built with) into a project-local `.venv` — no Java, no ANTLR
|
|
85
|
+
tooling required for that (those are dev-only, for regenerating
|
|
86
|
+
`src/metchurial/_generated/` or rebuilding `dist/metchurial.py` itself — see
|
|
87
|
+
[Dev workflow](#dev-workflow)).
|
|
88
|
+
|
|
89
|
+
### Installing from PyPI
|
|
90
|
+
|
|
91
|
+
```bash
|
|
92
|
+
pip install metchurial # or: uv add metchurial
|
|
93
|
+
```
|
|
94
|
+
|
|
95
|
+
This installs the `metchurial` CLI and the library API below. The
|
|
96
|
+
single-file `dist/metchurial.py` remains the distribution channel for
|
|
97
|
+
restricted environments where even `pip install` isn't an option.
|
|
98
|
+
|
|
99
|
+
### Using as a library
|
|
100
|
+
|
|
101
|
+
The CLI is one consumer of a plain Python API — install the package and
|
|
102
|
+
drive scans from your own code:
|
|
103
|
+
|
|
104
|
+
```python
|
|
105
|
+
import metchurial
|
|
106
|
+
|
|
107
|
+
# One call: scan a tree with every metadata analysis on.
|
|
108
|
+
result = metchurial.scan("/sql/root", metchurial.ScanOptions.metadata())
|
|
109
|
+
|
|
110
|
+
for f in result.findings: # sensitive-value findings
|
|
111
|
+
print(f.file, f.line, f.column_name, f.value)
|
|
112
|
+
for row in result.identity_rows: # per-statement core_ids
|
|
113
|
+
print(row.core_id, row.file, row.line)
|
|
114
|
+
|
|
115
|
+
# Or per-file, with only what you need switched on:
|
|
116
|
+
one = metchurial.scan_file(
|
|
117
|
+
"query.sql",
|
|
118
|
+
metchurial.ScanOptions(sensitive_columns=("ACCT_ID", "HLDR_NM"),
|
|
119
|
+
extract_relations=True))
|
|
120
|
+
```
|
|
121
|
+
|
|
122
|
+
`scan()`/`scan_file()` return typed result objects
|
|
123
|
+
(`TreeScanResult`/`FileScanResult` of `Finding`/`TableUse`/`RelationEdge`/
|
|
124
|
+
`IdentityRow`/... rows) and never print or write files on their own —
|
|
125
|
+
report artifacts (`summary.md`, `findings.tsv`, `refs_*.tsv`) are the
|
|
126
|
+
CLI's job. The one exception is `ScanOptions(split_selects=True)`, which
|
|
127
|
+
writes `-NN` split files next to each multi-SELECT source file, same as
|
|
128
|
+
the `--split-selects` flag. Everything a scan can be told is a field on
|
|
129
|
+
`ScanOptions` (a frozen dataclass); `ScanOptions.metadata()` is shorthand
|
|
130
|
+
for switching every `extract_*` analysis on, mirroring
|
|
131
|
+
`--extract-metadata`.
|
|
132
|
+
|
|
133
|
+
## CLI reference
|
|
134
|
+
|
|
135
|
+
| Flag | Default | Description |
|
|
136
|
+
|---|---|---|
|
|
137
|
+
| `root` (positional) | — | Directory to scan recursively |
|
|
138
|
+
| `--sensitive-columns` | `ACCT_ID CTRT_NO ACCT_NM ACCT_NAME` | Column names sensitive-column comparison detection treats as sensitive; fully replaces the default list, doesn't add to it |
|
|
139
|
+
| `--extensions` | `sql txt` | File extensions to scan, without the dot |
|
|
140
|
+
| `--extract-metadata` | off | Also emit `refs_tables.tsv`/`refs_columns.tsv`/`refs_functions.tsv`/`refs_relations.tsv`/`refs_query_identity.tsv` (schema/table/column refs, JOIN relationships, function/predicate usage, per-statement structural identity) and matching summary.md sections — see [Output artifacts](#output-artifacts) |
|
|
141
|
+
| `--query-similarity` | off | Also emit `refs_query_similarity.tsv`: pairwise Jaccard similarity between statements that don't share a `core_id`. Opt-in because the pass is O(n²) in the number of *distinct* core_ids — fine for thousands of distinct queries, slow for tens of thousands. Requires `--extract-metadata` |
|
|
142
|
+
| `--split-selects` | off | For a file with 2+ standalone SELECT blocks, write one `<stem>-NN<ext>` file per block alongside the original (files with a single block are left as-is) |
|
|
143
|
+
| `--mask-literals` | off | Rewrite in place every flagged literal's content to a fixed placeholder (`'****'`/`"****"` for quoted, `0000` for unquoted numeric), everything else byte-for-byte identical — back up files first, this overwrites them |
|
|
144
|
+
| `--workers N` | `1` | Scan across N worker processes instead of one |
|
|
145
|
+
| `--max-chunk-iterations N` | `200000` | Safety-valve cap on the resync driver's loop iterations per statement chunk |
|
|
146
|
+
| `--verbose` | off | Print a `[i/N]` progress line to stderr as each file is scanned |
|
|
147
|
+
|
|
148
|
+
`--workers N` scans files across N worker processes (`concurrent.futures.
|
|
149
|
+
ProcessPoolExecutor`, stdlib only). Parsing is CPU-bound pure Python, so
|
|
150
|
+
this is real multi-core parallelism, not threads, which the GIL would keep
|
|
151
|
+
from helping here. Each file is scanned independently with no shared
|
|
152
|
+
state, so results are unaffected other than which order they're merged in
|
|
153
|
+
(the reports already group/sort by file and line regardless). Leave a
|
|
154
|
+
couple of cores free for the OS/other work rather than setting `--workers`
|
|
155
|
+
to your full core count.
|
|
156
|
+
|
|
157
|
+
## What it extracts
|
|
158
|
+
|
|
159
|
+
`--extract-metadata` walks the same parse tree detection uses, but
|
|
160
|
+
unconditionally — every reference in the file, not just ones compared to a
|
|
161
|
+
literal — and resolves each one back to the table/schema it actually
|
|
162
|
+
belongs to:
|
|
163
|
+
|
|
164
|
+
- **Table & schema references** — every `schema.table` a file's SQL
|
|
165
|
+
touches. Each query block gets its own alias map, so a bare `t1` or
|
|
166
|
+
`t1.col` resolves back to the schema-qualified table it actually refers
|
|
167
|
+
to, not just the identifier as written on that line.
|
|
168
|
+
- **Column references** — every `schema.table.column` reference in the
|
|
169
|
+
file, not only ones inside a comparison; a correlated subquery's own
|
|
170
|
+
`t.col` is resolved within its own scope, not leaked into its parent's.
|
|
171
|
+
- **Function & predicate usage** — every function call (`SUBSTR`,
|
|
172
|
+
`COALESCE`, `SUM`, ...) and predicate operator (`=`, `IN`, `BETWEEN`,
|
|
173
|
+
`LIKE`, `IS NULL`, ...) actually used, with the exact source text of its
|
|
174
|
+
arguments/operands.
|
|
175
|
+
- **JOIN relationships** — every table-to-table JOIN edge, aggregated
|
|
176
|
+
across the *entire scan* (one graph, not one per file) — how tables in a
|
|
177
|
+
legacy schema actually connect in practice, not what an ER diagram
|
|
178
|
+
claims they should.
|
|
179
|
+
|
|
180
|
+
Each of these lands in its own `refs_*.tsv` — see
|
|
181
|
+
[Output artifacts](#output-artifacts) — built to be loaded straight into a
|
|
182
|
+
spreadsheet or a graph tool.
|
|
183
|
+
|
|
184
|
+
## What it detects
|
|
185
|
+
|
|
186
|
+
Sensitive-value detection is one specific analysis built on the same parse
|
|
187
|
+
tree metadata extraction uses — always on, independent of
|
|
188
|
+
`--extract-metadata`. It's two independent mechanisms, each producing its
|
|
189
|
+
own finding:
|
|
190
|
+
|
|
191
|
+
- **Sensitive-Column Comparison Detection — FINDING**: a sensitive column
|
|
192
|
+
(see `--sensitive-columns`) compared to a hardcoded literal — `=`, `<>`,
|
|
193
|
+
`!=`, `<=`, `>=`, `<`, `>`, `(NOT) IN (...)`, `(NOT) LIKE`,
|
|
194
|
+
`BETWEEN ... AND ...`, or a bare `(` before a literal (a DB2 quirk) — in
|
|
195
|
+
either direction and regardless of spacing or line breaks.
|
|
196
|
+
- **Known-Name Matching — FINDING**: any quoted, name-shaped literal (2-4
|
|
197
|
+
Hangul syllables) whose text is listed in `known_names.txt`, regardless
|
|
198
|
+
of which column it's compared to. There's no surname heuristic — a
|
|
199
|
+
literal only becomes a finding once a human has confirmed it's a real
|
|
200
|
+
name. Every other name-shaped literal is a triage *candidate*: it shows up
|
|
201
|
+
in `strings.txt` each run until it's copied into either `known_names.txt`
|
|
202
|
+
(flags it as a finding from then on) or `stopwords.txt` (excludes it from
|
|
203
|
+
`strings.txt` from then on). Both files are one word per line, `#`
|
|
204
|
+
comments allowed, auto-created empty on first run — see
|
|
205
|
+
[Output artifacts](#output-artifacts).
|
|
206
|
+
|
|
207
|
+
Findings inside `--`/`/* */` comments are still reported (commented-out
|
|
208
|
+
code can leak real data) and tagged `in_comment=Y` in `findings.tsv`.
|
|
209
|
+
`/* */` comments may nest, and a finding inside a nested comment is still
|
|
210
|
+
found.
|
|
211
|
+
|
|
212
|
+
### Public placeholder names
|
|
213
|
+
|
|
214
|
+
The column and table names used throughout this repo — `ACCT_ID`,
|
|
215
|
+
`CTRT_NO`, `ACCT_NM`, `ACCT_NAME`, `HLDR_NM`, `TBSAMPLE001`, `STAT_CD` —
|
|
216
|
+
are placeholders, not real production schema names from any actual DB2
|
|
217
|
+
environment, swapped in consistently before this repo was made public.
|
|
218
|
+
They appear in `DEFAULT_SENSITIVE_COLUMNS` (`src/metchurial/models/options.py`, `--sensitive-columns`'s
|
|
219
|
+
built-in default), every fixture under `tests/fixtures/`, and this
|
|
220
|
+
README's own examples. This doesn't affect behavior — `--sensitive-columns`
|
|
221
|
+
always fully replaces the default list, so a real deployment should pass
|
|
222
|
+
its own actual column names explicitly on every run rather than relying on
|
|
223
|
+
the shipped defaults meaning anything for your schema.
|
|
224
|
+
|
|
225
|
+
## Output artifacts
|
|
226
|
+
|
|
227
|
+
Every scan writes the same fixed set of files into the current working
|
|
228
|
+
directory (not the scan root, and not configurable — one predictable set
|
|
229
|
+
of names regardless of invocation). `summary.md` is an index into the
|
|
230
|
+
others: bounded counts and top-N tables with pointers to the full detail,
|
|
231
|
+
not a duplicate of it.
|
|
232
|
+
|
|
233
|
+
| File | Written when | Contents |
|
|
234
|
+
|---|---|---|
|
|
235
|
+
| `summary.md` | always | Run info, Sensitive Findings (with per-file detail), String Occurrences, Bad Files, Stopwords, Known Names, and — if enabled — Table & Column References, Functions, Relations, Select Blocks |
|
|
236
|
+
| `findings.tsv` | always | Every finding, one row per literal, for filtering/sorting in Excel |
|
|
237
|
+
| `strings.txt` | always | Unique name-shaped literals not yet classified into `known_names.txt`/`stopwords.txt`, with occurrence counts, in a format directly copy-pasteable into either |
|
|
238
|
+
| `stopwords.txt` | always (auto-created empty with a format header if missing) | Name-shaped literals reviewed and confirmed *not* sensitive — excluded from `strings.txt` from then on; edit in place |
|
|
239
|
+
| `known_names.txt` | always (auto-created empty with a format header if missing) | Name-shaped literals reviewed and confirmed sensitive — every matching literal becomes a known-name finding from then on; edit in place |
|
|
240
|
+
| `bad_files.txt` | always | Persistent skip-list of files too malformed to parse — see [Bad files](#bad-files) |
|
|
241
|
+
| `refs_tables.tsv` | `--extract-metadata` | Every `schema.table` reference found, with file/line |
|
|
242
|
+
| `refs_columns.tsv` | `--extract-metadata` | Every `schema.table.column` reference found, with file/line |
|
|
243
|
+
| `refs_functions.tsv` | `--extract-metadata` | Every function call and predicate operator found, with operands/file/line |
|
|
244
|
+
| `refs_relations.tsv` | `--extract-metadata` | Table-to-table JOIN usage aggregated across the whole scan (one file, not per-directory) |
|
|
245
|
+
| `refs_query_identity.tsv` | `--extract-metadata` | One `core_id` per statement — structurally identical statements share one id regardless of aliasing/projection/formatting differences |
|
|
246
|
+
| `refs_query_similarity.tsv` | `--query-similarity` | Pairwise Jaccard similarity between distinct `core_id`s that don't match exactly |
|
|
247
|
+
|
|
248
|
+
## Bad files
|
|
249
|
+
|
|
250
|
+
Some real-world SQL files aren't really valid SQL — internal section
|
|
251
|
+
dividers (`========`, `<<목표KPI>>`), bare prose headers, Korean-language
|
|
252
|
+
comments used as informal headings, or files with the actual SQL truncated
|
|
253
|
+
partway through. These can make the parser resync loop grind for a very
|
|
254
|
+
long time on a single file, or crash it outright.
|
|
255
|
+
|
|
256
|
+
Two independent safety nets guard against this:
|
|
257
|
+
|
|
258
|
+
- A cheap **pre-check** on the token stream (lexer-error ratio and
|
|
259
|
+
long runs of repeated punctuation) flags a file as bad before any real
|
|
260
|
+
parsing is attempted.
|
|
261
|
+
- A broad **try/except** around the actual scan of each file catches any
|
|
262
|
+
unexpected crash and treats it the same way.
|
|
263
|
+
|
|
264
|
+
Either path records the file's path and a short reason in `bad_files.txt`,
|
|
265
|
+
and the file is skipped entirely — not even attempted — on every later
|
|
266
|
+
run. Workflow:
|
|
267
|
+
|
|
268
|
+
1. Run the scan; anything unfixably weird lands in `bad_files.txt` and is
|
|
269
|
+
skipped from then on.
|
|
270
|
+
2. Fix the file's SQL content by hand (or decide it's fine to leave out).
|
|
271
|
+
3. Delete that file's line from `bad_files.txt`.
|
|
272
|
+
4. Re-run — the file is attempted again on the next scan.
|
|
273
|
+
|
|
274
|
+
`bad_files.txt` is a local, per-environment artifact (it's `.gitignore`d)
|
|
275
|
+
rather than something meant to be committed and shared.
|
|
276
|
+
|
|
277
|
+
## Known limitations
|
|
278
|
+
|
|
279
|
+
Known gaps in the vendored grammar, `--extract-metadata` extraction, and
|
|
280
|
+
the file-encoding auto-detection are tracked as GitHub issues, not
|
|
281
|
+
duplicated here — refer to
|
|
282
|
+
[Issues](https://github.com/CynicDog/metchurial/issues). Each is backed
|
|
283
|
+
by a runnable test (`tests/test_grammar_smoke.py`,
|
|
284
|
+
`tests/test_db2_grammar_specific_cases.py`, and friends), not just prose.
|
|
285
|
+
|
|
286
|
+
## How the bundle works
|
|
287
|
+
|
|
288
|
+
`dist/metchurial.py` is built with
|
|
289
|
+
[stickytape](https://github.com/mwilliamson/stickytape), which inlines
|
|
290
|
+
every module's source as embedded strings. **At every run**, it writes
|
|
291
|
+
that source out to a fresh OS temp directory, imports from there, and
|
|
292
|
+
deletes it on exit. Practically:
|
|
293
|
+
|
|
294
|
+
- It needs write access to the OS temp directory (normally fine, but worth
|
|
295
|
+
confirming in advance).
|
|
296
|
+
- Some corporate EDR/antivirus tools are wary of "a script writes many
|
|
297
|
+
`.py` files to temp and imports them" as a pattern. If your environment
|
|
298
|
+
has an infra security review step, flag this mechanism to them up front
|
|
299
|
+
— `docs/PROVENANCE.md` documents exactly what's bundled and why.
|
|
300
|
+
|
|
301
|
+
The file is large (~4.8MB, mostly the Db2 parser's serialized ATN tables)
|
|
302
|
+
-- one file, but not a small one, and not realistically human-auditable
|
|
303
|
+
top to bottom.
|
|
304
|
+
|
|
305
|
+
## Dev workflow
|
|
306
|
+
|
|
307
|
+
Everything below needs the dev dependency group (`antlr4-tools`,
|
|
308
|
+
`stickytape`, `ruff` — see `pyproject.toml`) plus Java for grammar
|
|
309
|
+
regeneration specifically (`uv`/`pip` can't install that) — none of this
|
|
310
|
+
runs in the restricted target environment; only `dist/metchurial.py` does.
|
|
311
|
+
|
|
312
|
+
```bash
|
|
313
|
+
uv sync
|
|
314
|
+
|
|
315
|
+
# regenerate src/metchurial/_generated/ from vendor/grammars-v4/*.g4 (only needed after
|
|
316
|
+
# touching the grammar itself)
|
|
317
|
+
uv run bash build/generate_parser.sh
|
|
318
|
+
|
|
319
|
+
# run everything: grammar smoke tests, scan_file()-level tests, end-to-end
|
|
320
|
+
# edge-case regressions, and the bundle self-containment check
|
|
321
|
+
uv run python -m unittest discover -s tests -p "test_*.py"
|
|
322
|
+
|
|
323
|
+
# lint
|
|
324
|
+
uv run ruff check src
|
|
325
|
+
|
|
326
|
+
# rebuild the deployable single-file artifact
|
|
327
|
+
uv run python build/bundle.py
|
|
328
|
+
|
|
329
|
+
# build the PyPI sdist+wheel (kept out of dist/, which holds the bundle)
|
|
330
|
+
uv build --out-dir pypi-dist
|
|
331
|
+
```
|
|
332
|
+
|
|
333
|
+
Publishing to PyPI is automated: publishing a GitHub release triggers
|
|
334
|
+
`.github/workflows/publish.yml`, which builds and uploads via PyPI
|
|
335
|
+
trusted publishing — no API tokens involved.
|
|
336
|
+
|
|
337
|
+
## Licensing
|
|
338
|
+
|
|
339
|
+
This project vendors two pieces of third-party code into the deployable
|
|
340
|
+
artifact: an IBM Db2 SQL grammar (MIT, from `antlr/grammars-v4`'s
|
|
341
|
+
`sql/db2`) and `antlr4-python3-runtime` (BSD-3-Clause). See
|
|
342
|
+
`docs/PROVENANCE.md` for exact versions, license texts, and what (if
|
|
343
|
+
anything) was modified.
|