py-tbparse 0.1.1a0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- py_tbparse-0.1.1a0/AGENTS.md +231 -0
- py_tbparse-0.1.1a0/LICENSE +25 -0
- py_tbparse-0.1.1a0/MANIFEST.in +6 -0
- py_tbparse-0.1.1a0/PKG-INFO +168 -0
- py_tbparse-0.1.1a0/README.md +135 -0
- py_tbparse-0.1.1a0/py_tbparse.egg-info/PKG-INFO +168 -0
- py_tbparse-0.1.1a0/py_tbparse.egg-info/SOURCES.txt +52 -0
- py_tbparse-0.1.1a0/py_tbparse.egg-info/dependency_links.txt +1 -0
- py_tbparse-0.1.1a0/py_tbparse.egg-info/entry_points.txt +3 -0
- py_tbparse-0.1.1a0/py_tbparse.egg-info/requires.txt +12 -0
- py_tbparse-0.1.1a0/py_tbparse.egg-info/top_level.txt +1 -0
- py_tbparse-0.1.1a0/pyproject.toml +49 -0
- py_tbparse-0.1.1a0/setup.cfg +4 -0
- py_tbparse-0.1.1a0/tests/conftest.py +25 -0
- py_tbparse-0.1.1a0/tests/fixtures/test_for_wenjie.twb +1899 -0
- py_tbparse-0.1.1a0/tests/fixtures/test_for_zip.twbx +0 -0
- py_tbparse-0.1.1a0/tests/test_batch.py +39 -0
- py_tbparse-0.1.1a0/tests/test_calculated_fields.py +20 -0
- py_tbparse-0.1.1a0/tests/test_clean.py +31 -0
- py_tbparse-0.1.1a0/tests/test_cli.py +137 -0
- py_tbparse-0.1.1a0/tests/test_dashboards.py +107 -0
- py_tbparse-0.1.1a0/tests/test_datasources.py +63 -0
- py_tbparse-0.1.1a0/tests/test_diff.py +68 -0
- py_tbparse-0.1.1a0/tests/test_fields.py +50 -0
- py_tbparse-0.1.1a0/tests/test_graph.py +70 -0
- py_tbparse-0.1.1a0/tests/test_gui_browser.py +228 -0
- py_tbparse-0.1.1a0/tests/test_joins.py +74 -0
- py_tbparse-0.1.1a0/tests/test_parser.py +81 -0
- py_tbparse-0.1.1a0/tests/test_published.py +52 -0
- py_tbparse-0.1.1a0/tests/test_relationships.py +62 -0
- py_tbparse-0.1.1a0/tests/test_sql.py +97 -0
- py_tbparse-0.1.1a0/tests/test_validators.py +62 -0
- py_tbparse-0.1.1a0/tests/test_webgui.py +228 -0
- py_tbparse-0.1.1a0/tests/test_xml.py +54 -0
- py_tbparse-0.1.1a0/twbparser_py/__init__.py +54 -0
- py_tbparse-0.1.1a0/twbparser_py/_clean.py +87 -0
- py_tbparse-0.1.1a0/twbparser_py/_tables.py +97 -0
- py_tbparse-0.1.1a0/twbparser_py/_xml.py +172 -0
- py_tbparse-0.1.1a0/twbparser_py/batch.py +62 -0
- py_tbparse-0.1.1a0/twbparser_py/calculated_fields.py +113 -0
- py_tbparse-0.1.1a0/twbparser_py/cli.py +207 -0
- py_tbparse-0.1.1a0/twbparser_py/dashboards.py +101 -0
- py_tbparse-0.1.1a0/twbparser_py/datasources.py +257 -0
- py_tbparse-0.1.1a0/twbparser_py/diff.py +86 -0
- py_tbparse-0.1.1a0/twbparser_py/fields.py +148 -0
- py_tbparse-0.1.1a0/twbparser_py/graph.py +81 -0
- py_tbparse-0.1.1a0/twbparser_py/joins.py +99 -0
- py_tbparse-0.1.1a0/twbparser_py/parser.py +234 -0
- py_tbparse-0.1.1a0/twbparser_py/published.py +52 -0
- py_tbparse-0.1.1a0/twbparser_py/py.typed +0 -0
- py_tbparse-0.1.1a0/twbparser_py/relationships.py +175 -0
- py_tbparse-0.1.1a0/twbparser_py/sql.py +76 -0
- py_tbparse-0.1.1a0/twbparser_py/validators.py +99 -0
- py_tbparse-0.1.1a0/twbparser_py/webgui.py +398 -0
|
@@ -0,0 +1,231 @@
|
|
|
1
|
+
# AGENTS.md
|
|
2
|
+
|
|
3
|
+
## What this is
|
|
4
|
+
|
|
5
|
+
`twbparser_py` is a native Python port of the R package
|
|
6
|
+
[`twbparser`](https://github.com/PrigasG/twbparser) (mirrored at
|
|
7
|
+
`DDSNA/twbparser`): it parses Tableau `.twb`/`.twbx` workbook files into
|
|
8
|
+
`pandas` DataFrames. Pure `lxml` XML parsing — no R runtime, no `rpy2`.
|
|
9
|
+
|
|
10
|
+
Current scope is a **v1 subset** of the original R package's ~50 exported
|
|
11
|
+
functions. Ported: workbook loading (`.twb`/`.twbx`), datasources,
|
|
12
|
+
parameters, raw/calculated fields, joins, relationships (legacy `<relation
|
|
13
|
+
type="join">` and 2020.2+ `<relationships>`), inferred relationships,
|
|
14
|
+
dashboards/dashboard-sheets, `validate_relationships`, custom/initial SQL,
|
|
15
|
+
and published-source detection. **Not yet ported** (v2): formatting,
|
|
16
|
+
tooltips, colors, axes, sorts, dashboard layout/actions, analytics
|
|
17
|
+
helpers (calc complexity, field usage, replication brief), and the
|
|
18
|
+
Shiny-inspector equivalent.
|
|
19
|
+
|
|
20
|
+
Also added, with no R equivalent (Python-native extras -- see "Non-R
|
|
21
|
+
modules" below): Graphviz DOT export of the relationship graph
|
|
22
|
+
(replaces the R package's igraph/ggraph-based plotting with a
|
|
23
|
+
dependency-free alternative), workbook-to-workbook diff, folder/batch
|
|
24
|
+
analysis across many workbooks, and Jupyter rich display.
|
|
25
|
+
|
|
26
|
+
## Setup
|
|
27
|
+
|
|
28
|
+
No system-wide `pip`; use the project venv.
|
|
29
|
+
|
|
30
|
+
```bash
|
|
31
|
+
python3 -m venv .venv
|
|
32
|
+
.venv/bin/pip install -e ".[test]"
|
|
33
|
+
```
|
|
34
|
+
|
|
35
|
+
## Build / test / lint commands
|
|
36
|
+
|
|
37
|
+
```bash
|
|
38
|
+
.venv/bin/python -m pytest -q # run the full test suite
|
|
39
|
+
.venv/bin/python -m pytest -q tests/test_joins.py # single file
|
|
40
|
+
.venv/bin/python -m py_compile twbparser_py/*.py # syntax check
|
|
41
|
+
```
|
|
42
|
+
|
|
43
|
+
Browser (GUI) tests need a one-time setup; without it they skip and the
|
|
44
|
+
rest of the suite still runs:
|
|
45
|
+
|
|
46
|
+
```bash
|
|
47
|
+
.venv/bin/pip install -e ".[browser]"
|
|
48
|
+
.venv/bin/playwright install chromium
|
|
49
|
+
./scripts/setup-browser-libs.sh # only on a machine without root
|
|
50
|
+
.venv/bin/python -m pytest -q tests/test_gui_browser.py
|
|
51
|
+
```
|
|
52
|
+
|
|
53
|
+
There is no linter/formatter configured yet — match existing style
|
|
54
|
+
(no trailing comments unless explaining non-obvious behavior, type hints
|
|
55
|
+
via `from __future__ import annotations`).
|
|
56
|
+
|
|
57
|
+
## Release / packaging
|
|
58
|
+
|
|
59
|
+
```bash
|
|
60
|
+
.venv/bin/pip install -e ".[dev]" # build + twine
|
|
61
|
+
rm -rf dist build twbparser_py.egg-info
|
|
62
|
+
.venv/bin/python -m build # produces dist/*.whl and dist/*.tar.gz
|
|
63
|
+
.venv/bin/twine check dist/* # validates metadata/README rendering
|
|
64
|
+
```
|
|
65
|
+
|
|
66
|
+
Version is single-sourced from `pyproject.toml`'s `[project].version`;
|
|
67
|
+
`twbparser_py.__version__` reads it back via `importlib.metadata` at
|
|
68
|
+
runtime (see `__init__.py`), so don't hardcode a second copy.
|
|
69
|
+
|
|
70
|
+
Before bumping the version for a release: bump `version` in
|
|
71
|
+
`pyproject.toml`, then rebuild and smoke-test the wheel in a
|
|
72
|
+
throwaway venv (`pip install dist/*.whl`, run `twbparser --help` and
|
|
73
|
+
`twbparser-gui --help`, run pytest against an extracted sdist) — this
|
|
74
|
+
catches packaging bugs (missing files, wrong entry points) that an
|
|
75
|
+
editable install won't.
|
|
76
|
+
|
|
77
|
+
CI (`.github/workflows/ci.yml`) runs pytest across Python 3.9–3.13, runs
|
|
78
|
+
the browser GUI tests in real Chromium (`gui` job), and builds+checks the
|
|
79
|
+
distribution on every push/PR. `build` depends on both `test` and `gui`
|
|
80
|
+
— don't drop `gui` from `build`'s `needs`, or a broken-page build can
|
|
81
|
+
succeed and publish an artifact while the one job that would have caught
|
|
82
|
+
it fails off to the side. Release
|
|
83
|
+
(`.github/workflows/release.yml`) publishes to PyPI via trusted
|
|
84
|
+
publishing (OIDC, no stored token) when a GitHub Release is published —
|
|
85
|
+
this needs to be configured once on the PyPI project's "Trusted
|
|
86
|
+
Publishers" settings page before the first release, and needs a
|
|
87
|
+
`release` GitHub Environment created in repo settings.
|
|
88
|
+
|
|
89
|
+
## Architecture
|
|
90
|
+
|
|
91
|
+
Each `twbparser_py/*.py` module is a direct port of one R source file in
|
|
92
|
+
the upstream package, function-for-function:
|
|
93
|
+
|
|
94
|
+
| Python module | Ported from (R) | Notes |
|
|
95
|
+
|---|---|---|
|
|
96
|
+
| `_clean.py` | `R/utils.R` (`.twb_clean_table`, `.twb_clean_field`, `attr_safe_get`, `.strip_brackets`) | Regex-based name cleaners shared by everything else |
|
|
97
|
+
| `_xml.py` | `R/utils.R` (`twbx_list`, `extract_twb_from_twbx`, `twbx_extract_files`) | `.twbx` zip handling |
|
|
98
|
+
| `datasources.py` | `R/datasources.R` | `extract_named_connections`, `extract_datasource_details`, `extract_parameters` |
|
|
99
|
+
| `fields.py` | `R/fields.R` | `extract_columns_with_table_source`, `infer_implicit_relationships` |
|
|
100
|
+
| `calculated_fields.py` | `R/calculated_fields.R` | `extract_calculated_fields`, `extract_raw_fields` |
|
|
101
|
+
| `joins.py` | `R/joins.R` | `extract_joins` (legacy `<relation type="join">`) |
|
|
102
|
+
| `relationships.py` | `R/relationships.R` | `extract_relations`, `build_object_table_mapping`, `extract_relationships` (2020.2+ model) |
|
|
103
|
+
| `dashboards.py` | `R/dashboard_details.R` (subset) | `list_dashboards`, `dashboard_sheets` |
|
|
104
|
+
| `validators.py` | `R/validators.R` | `validate_relationships` |
|
|
105
|
+
| `sql.py` | `R/sql.R` | `extract_custom_sql` (`twb_custom_sql`), `extract_initial_sql` (`twb_initial_sql`) |
|
|
106
|
+
| `published.py` | `R/published.R` | `extract_published_refs` (`twb_published_refs`) |
|
|
107
|
+
| `parser.py` | `R/twb_parser.R` (`TwbParser` R6 class) | The `TwbParser` façade tying every module together |
|
|
108
|
+
|
|
109
|
+
`parser.py`'s `TwbParser.__init__` mirrors the R6 constructor: it eagerly
|
|
110
|
+
computes every cached DataFrame, guarding each with `_safe_call` (the port
|
|
111
|
+
of R's `safe_call`/`tryCatch`) so a malformed workbook degrades to empty,
|
|
112
|
+
correctly-columned DataFrames instead of raising.
|
|
113
|
+
|
|
114
|
+
### Non-R modules
|
|
115
|
+
|
|
116
|
+
These have no upstream R source to track — they're either the
|
|
117
|
+
Python-native user-facing layer on top of `TwbParser`, or features added
|
|
118
|
+
beyond the R package's scope:
|
|
119
|
+
|
|
120
|
+
| Python module | What it is |
|
|
121
|
+
|---|---|
|
|
122
|
+
| `_tables.py` | Name → `TwbParser`-accessor registry shared by `cli.py`, `webgui.py`, `diff.py`, and `batch.py`. Adding a new extractor to `TwbParser`? Add it here too so it's automatically available everywhere else. |
|
|
123
|
+
| `cli.py` | `twbparser` command-line entry point, plus the `diff`/`batch` subcommands (dispatched on `sys.argv[1]` before the normal single-workbook argparse parser runs) |
|
|
124
|
+
| `webgui.py` | `twbparser-gui`: stdlib-only (`http.server` + vanilla JS) local browser GUI, no GUI toolkit dependency. Loosely fills the role of the R package's `run_twbparser_app`/Shiny inspector. |
|
|
125
|
+
| `graph.py` | `to_dot()`: Graphviz DOT export of joins/relationships (+ optional inferred, as dashed edges). Replaces the R package's igraph/ggraph-based `plot_dependency_graph`/`plot_relationship_graph` with a dependency-free text format any Graphviz-compatible tool can render. |
|
|
126
|
+
| `diff.py` | `diff_tables()`/`diff_workbooks()`: row-level added/removed diff between two workbooks' same-named table, via `_tables.TABLE_SPECS`. No "changed" classification without a natural key — a changed row shows as one removed + one added row. |
|
|
127
|
+
| `batch.py` | `scan_folder()`: runs one table across every `.twb`/`.twbx` in a directory, concatenated with a `workbook` column. Skips (with a warning) any file that fails to load/extract rather than aborting the batch. |
|
|
128
|
+
|
|
129
|
+
## Porting conventions (read before adding/modifying a function)
|
|
130
|
+
|
|
131
|
+
1. **Go to the R source first.** Fetch the corresponding file from
|
|
132
|
+
`github.com/DDSNA/twbparser` (or `PrigasG/twbparser`, same content) via
|
|
133
|
+
the GitHub API/raw URLs — don't guess at Tableau XML schema from
|
|
134
|
+
memory. Port the XPath expressions and fallback logic as literally as
|
|
135
|
+
possible; deviations from the R behavior are bugs, not improvements.
|
|
136
|
+
2. **XPath parity**: `lxml`'s XPath 1.0 engine supports the same
|
|
137
|
+
functions R's `xml2`/libxml2 use (`local-name()`, `contains()`,
|
|
138
|
+
`count()`), so most R XPath strings can be copied verbatim into
|
|
139
|
+
`.xpath(...)` calls.
|
|
140
|
+
3. **Empty-input contract**: every extractor returns an empty DataFrame
|
|
141
|
+
with the *correct columns* (not just `pd.DataFrame()`) when there's no
|
|
142
|
+
matching XML — callers (esp. `parser.py`) rely on this. Each module
|
|
143
|
+
exposes its column list as a `_XXX_COLUMNS` constant (e.g.
|
|
144
|
+
`joins._JOIN_COLUMNS`, `datasources._DATASOURCE_COLUMNS`) precisely so
|
|
145
|
+
`parser.py`'s `_safe_call` fallbacks can reuse it instead of retyping
|
|
146
|
+
the list a second time (which drifted out of sync once already).
|
|
147
|
+
4. **Cleaning helpers**: always route table/field name cleanup through
|
|
148
|
+
`_clean.clean_table` / `_clean.clean_field` / `_clean.strip_brackets`
|
|
149
|
+
rather than re-deriving regexes inline. Same for "is this value
|
|
150
|
+
missing" checks — use `_clean.is_missing(x)`, not a hand-rolled
|
|
151
|
+
`pd.isna()`/`isinstance(x, float)` check (its edge cases, like NaT or
|
|
152
|
+
array-likes, are easy to get subtly wrong per call site).
|
|
153
|
+
5. **No R dependency, ever.** If a future port needs something R gets
|
|
154
|
+
from `dplyr`/`igraph` for free (e.g. graph layout for
|
|
155
|
+
`plot_dependency_graph`), find a pure-Python equivalent or scope it
|
|
156
|
+
out — don't reach for `rpy2`/subprocess-to-R.
|
|
157
|
+
6. **XPath string literals**: never build one with
|
|
158
|
+
`f"...='{value}'"` or a `.replace("'", "")` — user/workbook-controlled
|
|
159
|
+
values (a dashboard name, say) can contain `'`, `"`, or both, and
|
|
160
|
+
XPath 1.0 has no in-literal escape character. Use
|
|
161
|
+
`dashboards._xpath_string_literal(value)` (or extend it if a new
|
|
162
|
+
module needs the same thing), which picks a quoting style or falls
|
|
163
|
+
back to `concat()` as needed — verified against real `lxml` evaluation
|
|
164
|
+
for the tricky cases (leading/trailing/consecutive quotes, both quote
|
|
165
|
+
types at once).
|
|
166
|
+
|
|
167
|
+
## Testing conventions
|
|
168
|
+
|
|
169
|
+
- Fixtures live in `tests/fixtures/` (`test_for_wenjie.twb`,
|
|
170
|
+
`test_for_zip.twbx`) — pulled directly from the R package's
|
|
171
|
+
`inst/extdata/`. Don't regenerate/hand-edit them; if a new fixture is
|
|
172
|
+
needed, pull it from upstream the same way, or build a minimal
|
|
173
|
+
synthetic XML snippet (see `tests/test_joins.py`,
|
|
174
|
+
`tests/test_dashboards.py` for the pattern — many are lifted from the R
|
|
175
|
+
functions' own `@examples` roxygen blocks).
|
|
176
|
+
- `tests/conftest.py` provides `wenjie_xml`, `wenjie_path`,
|
|
177
|
+
`zip_twbx_path` fixtures and an `xml_from_string()` helper. Tests import
|
|
178
|
+
it with `from conftest import xml_from_string` (no `tests/__init__.py`,
|
|
179
|
+
so it's a plain top-level import, not a relative one).
|
|
180
|
+
- Every new extractor function needs: one test against a real fixture (if
|
|
181
|
+
it exercises fixture data) and/or one synthetic-XML test for edge
|
|
182
|
+
cases the fixtures don't cover (empty input, fallback paths).
|
|
183
|
+
- When in doubt about expected output, that's a signal to install R +
|
|
184
|
+
the `twbparser` package and diff against the real function's output on
|
|
185
|
+
the same fixture, rather than guessing.
|
|
186
|
+
|
|
187
|
+
### Testing the GUI
|
|
188
|
+
|
|
189
|
+
`tests/test_webgui.py` drives the HTTP endpoints directly — it never
|
|
190
|
+
executes the page's JavaScript. That blind spot shipped a real bug: a
|
|
191
|
+
`SyntaxError` in the inline `<script>` (a `\n` in a non-raw Python
|
|
192
|
+
string became a literal newline, splitting a JS string literal across
|
|
193
|
+
two lines) killed the entire script, so no handlers bound and the UI was
|
|
194
|
+
inert — with the whole suite green.
|
|
195
|
+
|
|
196
|
+
So: **any change to `webgui.py`'s `_PAGE` needs a browser test**, in
|
|
197
|
+
`tests/test_gui_browser.py`, which runs the page in real headless
|
|
198
|
+
Chromium via Playwright and fails on any uncaught JS error. The cheap
|
|
199
|
+
structural guards in `test_webgui.py` (unterminated string literals,
|
|
200
|
+
bracket balance) are a backstop, not a substitute.
|
|
201
|
+
|
|
202
|
+
`_PAGE` is declared as `r"""..."""` (a **raw** string) specifically so
|
|
203
|
+
this can't recur: without `r`, any backslash escape meant for the
|
|
204
|
+
*browser* (`\n`, `\t`, a future `\'`) would need doubling in the Python
|
|
205
|
+
source, and forgetting to double it silently corrupts the embedded JS
|
|
206
|
+
instead of erroring. Keep it raw — write JS escapes the normal JS way
|
|
207
|
+
(`'\n'`, not `'\\n'`).
|
|
208
|
+
|
|
209
|
+
Any server-side value spliced into `_PAGE` (currently `TABLE_NAMES` and
|
|
210
|
+
the preloaded workbook path) must go through `_json_for_script()`, not
|
|
211
|
+
bare `json.dumps()`. `json.dumps` doesn't escape `/`, so a value
|
|
212
|
+
containing the literal text `</script>` closes the script tag early in
|
|
213
|
+
the browser's HTML parser — this is real, not theoretical, since the
|
|
214
|
+
workbook path is user-controlled input; see
|
|
215
|
+
`test_preload_path_cannot_break_out_of_script_tag` for a reproduction.
|
|
216
|
+
|
|
217
|
+
`scripts/setup-browser-libs.sh` unpacks Chromium's system libraries into
|
|
218
|
+
a gitignored `.browser-libs/` instead of apt-installing them as root;
|
|
219
|
+
the test fixture picks that directory up automatically via
|
|
220
|
+
`LD_LIBRARY_PATH`. It verifies each download against the checksum
|
|
221
|
+
`apt-get --print-uris` reports for it (SHA256 if offered, MD5Sum as the
|
|
222
|
+
realistic fallback — this environment's apt only emits MD5Sum) before
|
|
223
|
+
extracting; don't remove that step to "simplify" the script. Don't add
|
|
224
|
+
`pytest-playwright` — it's a pytest plugin that imports playwright at
|
|
225
|
+
startup, which makes collection fail for anyone who doesn't have it
|
|
226
|
+
installed.
|
|
227
|
+
|
|
228
|
+
## Commit / PR conventions
|
|
229
|
+
|
|
230
|
+
Nothing project-specific beyond the harness defaults — see repo commit
|
|
231
|
+
history for message style.
|
|
@@ -0,0 +1,25 @@
|
|
|
1
|
+
MIT License
|
|
2
|
+
|
|
3
|
+
This project is a Python port of logic originally implemented in the R
|
|
4
|
+
package "twbparser" (https://github.com/PrigasG/twbparser),
|
|
5
|
+
Copyright (c) 2025 George Arthur.
|
|
6
|
+
|
|
7
|
+
Copyright (c) 2026 the twbparser-py contributors
|
|
8
|
+
|
|
9
|
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
|
10
|
+
of this software and associated documentation files (the "Software"), to deal
|
|
11
|
+
in the Software without restriction, including without limitation the rights
|
|
12
|
+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
|
13
|
+
copies of the Software, and to permit persons to whom the Software is
|
|
14
|
+
furnished to do so, subject to the following conditions:
|
|
15
|
+
|
|
16
|
+
The above copyright notice and this permission notice shall be included in all
|
|
17
|
+
copies or substantial portions of the Software.
|
|
18
|
+
|
|
19
|
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
|
20
|
+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
|
21
|
+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
|
22
|
+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
|
23
|
+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
|
24
|
+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
|
25
|
+
SOFTWARE.
|
|
@@ -0,0 +1,168 @@
|
|
|
1
|
+
Metadata-Version: 2.4
|
|
2
|
+
Name: py-tbparse
|
|
3
|
+
Version: 0.1.1a0
|
|
4
|
+
Summary: Native Python port of the twbparser R package: parse Tableau .twb/.twbx workbooks into pandas DataFrames.
|
|
5
|
+
Author: DDSNA
|
|
6
|
+
License: MIT
|
|
7
|
+
Keywords: tableau,twb,twbx,workbook,parser,pandas
|
|
8
|
+
Classifier: Development Status :: 4 - Beta
|
|
9
|
+
Classifier: Intended Audience :: Developers
|
|
10
|
+
Classifier: License :: OSI Approved :: MIT License
|
|
11
|
+
Classifier: Operating System :: OS Independent
|
|
12
|
+
Classifier: Programming Language :: Python :: 3
|
|
13
|
+
Classifier: Programming Language :: Python :: 3.9
|
|
14
|
+
Classifier: Programming Language :: Python :: 3.10
|
|
15
|
+
Classifier: Programming Language :: Python :: 3.11
|
|
16
|
+
Classifier: Programming Language :: Python :: 3.12
|
|
17
|
+
Classifier: Programming Language :: Python :: 3.13
|
|
18
|
+
Classifier: Topic :: Software Development :: Libraries :: Python Modules
|
|
19
|
+
Classifier: Topic :: Office/Business
|
|
20
|
+
Requires-Python: >=3.9
|
|
21
|
+
Description-Content-Type: text/markdown
|
|
22
|
+
License-File: LICENSE
|
|
23
|
+
Requires-Dist: lxml>=4.9
|
|
24
|
+
Requires-Dist: pandas>=1.5
|
|
25
|
+
Provides-Extra: test
|
|
26
|
+
Requires-Dist: pytest>=7; extra == "test"
|
|
27
|
+
Provides-Extra: browser
|
|
28
|
+
Requires-Dist: playwright>=1.40; extra == "browser"
|
|
29
|
+
Provides-Extra: dev
|
|
30
|
+
Requires-Dist: build>=1.0; extra == "dev"
|
|
31
|
+
Requires-Dist: twine>=5.0; extra == "dev"
|
|
32
|
+
Dynamic: license-file
|
|
33
|
+
|
|
34
|
+
# twbparser-py
|
|
35
|
+
|
|
36
|
+
A native Python port of the [`twbparser`](https://github.com/PrigasG/twbparser)
|
|
37
|
+
R package: parses Tableau `.twb`/`.twbx` workbook files into `pandas`
|
|
38
|
+
DataFrames. No R runtime required — pure `lxml` XML parsing.
|
|
39
|
+
|
|
40
|
+
This is a v1 subset covering the parser's core: workbook loading,
|
|
41
|
+
datasources, parameters, fields, calculated fields, joins, relationships
|
|
42
|
+
(legacy and 2020.2+), inferred relationships, dashboards, relationship
|
|
43
|
+
validation, custom/initial SQL, and published-source detection.
|
|
44
|
+
Formatting/tooltips/colors/axes/sorts, dashboard layout/actions,
|
|
45
|
+
analytics helpers (calc complexity, field usage, replication brief), and
|
|
46
|
+
the Shiny-inspector equivalent are not yet ported.
|
|
47
|
+
|
|
48
|
+
Beyond the R original, this port also adds a few Python-native extras: a
|
|
49
|
+
Graphviz DOT export of the relationship graph, a workbook-to-workbook
|
|
50
|
+
diff, folder/batch analysis across many workbooks, and Jupyter rich
|
|
51
|
+
display (`_repr_html_`).
|
|
52
|
+
|
|
53
|
+
## Install
|
|
54
|
+
|
|
55
|
+
```bash
|
|
56
|
+
pip install -e .
|
|
57
|
+
```
|
|
58
|
+
|
|
59
|
+
## Usage
|
|
60
|
+
|
|
61
|
+
```python
|
|
62
|
+
from twbparser_py import TwbParser
|
|
63
|
+
|
|
64
|
+
p = TwbParser("workbook.twb") # or .twbx
|
|
65
|
+
p.get_datasources()
|
|
66
|
+
p.get_fields()
|
|
67
|
+
p.get_calculated_fields()
|
|
68
|
+
p.get_joins()
|
|
69
|
+
p.get_relationships()
|
|
70
|
+
p.get_inferred_relationships()
|
|
71
|
+
p.get_dashboards()
|
|
72
|
+
p.get_dashboard_sheets()
|
|
73
|
+
p.get_custom_sql()
|
|
74
|
+
p.get_initial_sql()
|
|
75
|
+
p.get_published_refs()
|
|
76
|
+
p.get_relationship_graph_dot() # Graphviz DOT string
|
|
77
|
+
p.validate()
|
|
78
|
+
p.get_overview()
|
|
79
|
+
p # in Jupyter: renders get_overview() via _repr_html_
|
|
80
|
+
|
|
81
|
+
from twbparser_py import diff_workbooks, scan_folder
|
|
82
|
+
|
|
83
|
+
diff_workbooks(TwbParser("v1.twb"), TwbParser("v2.twb"), table="datasources")
|
|
84
|
+
scan_folder("./workbooks", table="datasources") # one row per workbook x datasource
|
|
85
|
+
```
|
|
86
|
+
|
|
87
|
+
### CLI
|
|
88
|
+
|
|
89
|
+
```bash
|
|
90
|
+
twbparser workbook.twb # overview (default table)
|
|
91
|
+
twbparser workbook.twb tables # list available tables
|
|
92
|
+
twbparser workbook.twb calculated-fields # print a table
|
|
93
|
+
twbparser workbook.twb fields --format csv -o fields.csv
|
|
94
|
+
twbparser workbook.twbx dashboard-sheets --dashboard "Sales Overview"
|
|
95
|
+
twbparser workbook.twb validate # exit code 2 if invalid
|
|
96
|
+
twbparser workbook.twb graph --include-inferred > relationships.dot
|
|
97
|
+
twbparser diff old.twb new.twb datasources # row-level added/removed
|
|
98
|
+
twbparser batch ./workbooks datasources # one table, every workbook in a folder
|
|
99
|
+
```
|
|
100
|
+
|
|
101
|
+
Tables: `overview`, `datasources`, `parameters`, `fields`, `raw-fields`,
|
|
102
|
+
`calculated-fields`, `joins`, `relations`, `relationships`,
|
|
103
|
+
`inferred-relationships`, `dashboards`, `dashboard-sheets`,
|
|
104
|
+
`custom-sql`, `initial-sql`, `published-refs`. `--format` is `table`
|
|
105
|
+
(default), `csv`, or `json` (`graph` always prints Graphviz DOT text
|
|
106
|
+
regardless of `--format`). `diff`/`batch` accept most of the same table
|
|
107
|
+
names, minus `graph`/`validate`/`tables`.
|
|
108
|
+
|
|
109
|
+
### GUI
|
|
110
|
+
|
|
111
|
+
A local, browser-based GUI — standard library only (`http.server` +
|
|
112
|
+
vanilla JS), no GUI toolkit or extra dependency required:
|
|
113
|
+
|
|
114
|
+
```bash
|
|
115
|
+
twbparser-gui workbook.twb # opens your default browser
|
|
116
|
+
twbparser-gui # opens with an empty path field; paste one and click Load
|
|
117
|
+
twbparser-gui --no-browser --port 8765 # just run the server, e.g. for a headless box
|
|
118
|
+
```
|
|
119
|
+
|
|
120
|
+
Pick a table from the dropdown, filter `dashboard-sheets` by dashboard,
|
|
121
|
+
toggle "include Parameters" for `calculated-fields`, pick `graph` to
|
|
122
|
+
preview/export a Graphviz DOT digraph of the relationships, and export
|
|
123
|
+
any tabular view as CSV. All state lives server-side in memory for the
|
|
124
|
+
life of the process — it's a single-user local tool, not something to
|
|
125
|
+
expose on a shared network.
|
|
126
|
+
|
|
127
|
+
## Testing
|
|
128
|
+
|
|
129
|
+
```bash
|
|
130
|
+
pip install -e ".[test]"
|
|
131
|
+
pytest
|
|
132
|
+
```
|
|
133
|
+
|
|
134
|
+
Fixtures in `tests/fixtures/` are the same tiny sample workbooks used by
|
|
135
|
+
the original R package's test suite (`inst/extdata/`).
|
|
136
|
+
|
|
137
|
+
The GUI additionally has end-to-end tests that drive the page in a real
|
|
138
|
+
headless Chromium (via Playwright) and fail on any uncaught JavaScript
|
|
139
|
+
error. They're opt-in — without the browser installed they skip and the
|
|
140
|
+
rest of the suite runs normally:
|
|
141
|
+
|
|
142
|
+
```bash
|
|
143
|
+
pip install -e ".[test,browser]"
|
|
144
|
+
playwright install chromium # add --with-deps if you have root
|
|
145
|
+
./scripts/setup-browser-libs.sh # no-root alternative to --with-deps
|
|
146
|
+
pytest tests/test_gui_browser.py
|
|
147
|
+
```
|
|
148
|
+
|
|
149
|
+
## Background
|
|
150
|
+
|
|
151
|
+
`.twb` is plain XML, and `.twbx` is just a zip wrapper around one, which
|
|
152
|
+
is why parsing it natively in Python — no R, no reverse-engineering —
|
|
153
|
+
was tractable at this scale. Power BI's equivalent format, `.pbix`, is a
|
|
154
|
+
binary container built around the proprietary VertiPaq storage engine,
|
|
155
|
+
which is why reading it programmatically needed dedicated
|
|
156
|
+
reverse-engineering projects like
|
|
157
|
+
[PBIXRay](https://github.com/Hugoberry/pbixray) and
|
|
158
|
+
[pbi-tools](https://github.com/pbi-tools/pbi-tools). The comparison
|
|
159
|
+
isn't one-sided, though: Tableau's own official Python tooling for
|
|
160
|
+
*server* automation
|
|
161
|
+
([`tableauserverclient`](https://pypi.org/project/tableauserverclient/),
|
|
162
|
+
`tabcmd`) is more mature and more open than anything Microsoft ships for
|
|
163
|
+
Power BI's REST API.
|
|
164
|
+
|
|
165
|
+
## Credit
|
|
166
|
+
|
|
167
|
+
Ported from the R implementation by George Arthur
|
|
168
|
+
([`PrigasG/twbparser`](https://github.com/PrigasG/twbparser)), MIT licensed.
|
|
@@ -0,0 +1,135 @@
|
|
|
1
|
+
# twbparser-py
|
|
2
|
+
|
|
3
|
+
A native Python port of the [`twbparser`](https://github.com/PrigasG/twbparser)
|
|
4
|
+
R package: parses Tableau `.twb`/`.twbx` workbook files into `pandas`
|
|
5
|
+
DataFrames. No R runtime required — pure `lxml` XML parsing.
|
|
6
|
+
|
|
7
|
+
This is a v1 subset covering the parser's core: workbook loading,
|
|
8
|
+
datasources, parameters, fields, calculated fields, joins, relationships
|
|
9
|
+
(legacy and 2020.2+), inferred relationships, dashboards, relationship
|
|
10
|
+
validation, custom/initial SQL, and published-source detection.
|
|
11
|
+
Formatting/tooltips/colors/axes/sorts, dashboard layout/actions,
|
|
12
|
+
analytics helpers (calc complexity, field usage, replication brief), and
|
|
13
|
+
the Shiny-inspector equivalent are not yet ported.
|
|
14
|
+
|
|
15
|
+
Beyond the R original, this port also adds a few Python-native extras: a
|
|
16
|
+
Graphviz DOT export of the relationship graph, a workbook-to-workbook
|
|
17
|
+
diff, folder/batch analysis across many workbooks, and Jupyter rich
|
|
18
|
+
display (`_repr_html_`).
|
|
19
|
+
|
|
20
|
+
## Install
|
|
21
|
+
|
|
22
|
+
```bash
|
|
23
|
+
pip install -e .
|
|
24
|
+
```
|
|
25
|
+
|
|
26
|
+
## Usage
|
|
27
|
+
|
|
28
|
+
```python
|
|
29
|
+
from twbparser_py import TwbParser
|
|
30
|
+
|
|
31
|
+
p = TwbParser("workbook.twb") # or .twbx
|
|
32
|
+
p.get_datasources()
|
|
33
|
+
p.get_fields()
|
|
34
|
+
p.get_calculated_fields()
|
|
35
|
+
p.get_joins()
|
|
36
|
+
p.get_relationships()
|
|
37
|
+
p.get_inferred_relationships()
|
|
38
|
+
p.get_dashboards()
|
|
39
|
+
p.get_dashboard_sheets()
|
|
40
|
+
p.get_custom_sql()
|
|
41
|
+
p.get_initial_sql()
|
|
42
|
+
p.get_published_refs()
|
|
43
|
+
p.get_relationship_graph_dot() # Graphviz DOT string
|
|
44
|
+
p.validate()
|
|
45
|
+
p.get_overview()
|
|
46
|
+
p # in Jupyter: renders get_overview() via _repr_html_
|
|
47
|
+
|
|
48
|
+
from twbparser_py import diff_workbooks, scan_folder
|
|
49
|
+
|
|
50
|
+
diff_workbooks(TwbParser("v1.twb"), TwbParser("v2.twb"), table="datasources")
|
|
51
|
+
scan_folder("./workbooks", table="datasources") # one row per workbook x datasource
|
|
52
|
+
```
|
|
53
|
+
|
|
54
|
+
### CLI
|
|
55
|
+
|
|
56
|
+
```bash
|
|
57
|
+
twbparser workbook.twb # overview (default table)
|
|
58
|
+
twbparser workbook.twb tables # list available tables
|
|
59
|
+
twbparser workbook.twb calculated-fields # print a table
|
|
60
|
+
twbparser workbook.twb fields --format csv -o fields.csv
|
|
61
|
+
twbparser workbook.twbx dashboard-sheets --dashboard "Sales Overview"
|
|
62
|
+
twbparser workbook.twb validate # exit code 2 if invalid
|
|
63
|
+
twbparser workbook.twb graph --include-inferred > relationships.dot
|
|
64
|
+
twbparser diff old.twb new.twb datasources # row-level added/removed
|
|
65
|
+
twbparser batch ./workbooks datasources # one table, every workbook in a folder
|
|
66
|
+
```
|
|
67
|
+
|
|
68
|
+
Tables: `overview`, `datasources`, `parameters`, `fields`, `raw-fields`,
|
|
69
|
+
`calculated-fields`, `joins`, `relations`, `relationships`,
|
|
70
|
+
`inferred-relationships`, `dashboards`, `dashboard-sheets`,
|
|
71
|
+
`custom-sql`, `initial-sql`, `published-refs`. `--format` is `table`
|
|
72
|
+
(default), `csv`, or `json` (`graph` always prints Graphviz DOT text
|
|
73
|
+
regardless of `--format`). `diff`/`batch` accept most of the same table
|
|
74
|
+
names, minus `graph`/`validate`/`tables`.
|
|
75
|
+
|
|
76
|
+
### GUI
|
|
77
|
+
|
|
78
|
+
A local, browser-based GUI — standard library only (`http.server` +
|
|
79
|
+
vanilla JS), no GUI toolkit or extra dependency required:
|
|
80
|
+
|
|
81
|
+
```bash
|
|
82
|
+
twbparser-gui workbook.twb # opens your default browser
|
|
83
|
+
twbparser-gui # opens with an empty path field; paste one and click Load
|
|
84
|
+
twbparser-gui --no-browser --port 8765 # just run the server, e.g. for a headless box
|
|
85
|
+
```
|
|
86
|
+
|
|
87
|
+
Pick a table from the dropdown, filter `dashboard-sheets` by dashboard,
|
|
88
|
+
toggle "include Parameters" for `calculated-fields`, pick `graph` to
|
|
89
|
+
preview/export a Graphviz DOT digraph of the relationships, and export
|
|
90
|
+
any tabular view as CSV. All state lives server-side in memory for the
|
|
91
|
+
life of the process — it's a single-user local tool, not something to
|
|
92
|
+
expose on a shared network.
|
|
93
|
+
|
|
94
|
+
## Testing
|
|
95
|
+
|
|
96
|
+
```bash
|
|
97
|
+
pip install -e ".[test]"
|
|
98
|
+
pytest
|
|
99
|
+
```
|
|
100
|
+
|
|
101
|
+
Fixtures in `tests/fixtures/` are the same tiny sample workbooks used by
|
|
102
|
+
the original R package's test suite (`inst/extdata/`).
|
|
103
|
+
|
|
104
|
+
The GUI additionally has end-to-end tests that drive the page in a real
|
|
105
|
+
headless Chromium (via Playwright) and fail on any uncaught JavaScript
|
|
106
|
+
error. They're opt-in — without the browser installed they skip and the
|
|
107
|
+
rest of the suite runs normally:
|
|
108
|
+
|
|
109
|
+
```bash
|
|
110
|
+
pip install -e ".[test,browser]"
|
|
111
|
+
playwright install chromium # add --with-deps if you have root
|
|
112
|
+
./scripts/setup-browser-libs.sh # no-root alternative to --with-deps
|
|
113
|
+
pytest tests/test_gui_browser.py
|
|
114
|
+
```
|
|
115
|
+
|
|
116
|
+
## Background
|
|
117
|
+
|
|
118
|
+
`.twb` is plain XML, and `.twbx` is just a zip wrapper around one, which
|
|
119
|
+
is why parsing it natively in Python — no R, no reverse-engineering —
|
|
120
|
+
was tractable at this scale. Power BI's equivalent format, `.pbix`, is a
|
|
121
|
+
binary container built around the proprietary VertiPaq storage engine,
|
|
122
|
+
which is why reading it programmatically needed dedicated
|
|
123
|
+
reverse-engineering projects like
|
|
124
|
+
[PBIXRay](https://github.com/Hugoberry/pbixray) and
|
|
125
|
+
[pbi-tools](https://github.com/pbi-tools/pbi-tools). The comparison
|
|
126
|
+
isn't one-sided, though: Tableau's own official Python tooling for
|
|
127
|
+
*server* automation
|
|
128
|
+
([`tableauserverclient`](https://pypi.org/project/tableauserverclient/),
|
|
129
|
+
`tabcmd`) is more mature and more open than anything Microsoft ships for
|
|
130
|
+
Power BI's REST API.
|
|
131
|
+
|
|
132
|
+
## Credit
|
|
133
|
+
|
|
134
|
+
Ported from the R implementation by George Arthur
|
|
135
|
+
([`PrigasG/twbparser`](https://github.com/PrigasG/twbparser)), MIT licensed.
|