py-tbparse 0.1.1a0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (54) hide show
  1. py_tbparse-0.1.1a0/AGENTS.md +231 -0
  2. py_tbparse-0.1.1a0/LICENSE +25 -0
  3. py_tbparse-0.1.1a0/MANIFEST.in +6 -0
  4. py_tbparse-0.1.1a0/PKG-INFO +168 -0
  5. py_tbparse-0.1.1a0/README.md +135 -0
  6. py_tbparse-0.1.1a0/py_tbparse.egg-info/PKG-INFO +168 -0
  7. py_tbparse-0.1.1a0/py_tbparse.egg-info/SOURCES.txt +52 -0
  8. py_tbparse-0.1.1a0/py_tbparse.egg-info/dependency_links.txt +1 -0
  9. py_tbparse-0.1.1a0/py_tbparse.egg-info/entry_points.txt +3 -0
  10. py_tbparse-0.1.1a0/py_tbparse.egg-info/requires.txt +12 -0
  11. py_tbparse-0.1.1a0/py_tbparse.egg-info/top_level.txt +1 -0
  12. py_tbparse-0.1.1a0/pyproject.toml +49 -0
  13. py_tbparse-0.1.1a0/setup.cfg +4 -0
  14. py_tbparse-0.1.1a0/tests/conftest.py +25 -0
  15. py_tbparse-0.1.1a0/tests/fixtures/test_for_wenjie.twb +1899 -0
  16. py_tbparse-0.1.1a0/tests/fixtures/test_for_zip.twbx +0 -0
  17. py_tbparse-0.1.1a0/tests/test_batch.py +39 -0
  18. py_tbparse-0.1.1a0/tests/test_calculated_fields.py +20 -0
  19. py_tbparse-0.1.1a0/tests/test_clean.py +31 -0
  20. py_tbparse-0.1.1a0/tests/test_cli.py +137 -0
  21. py_tbparse-0.1.1a0/tests/test_dashboards.py +107 -0
  22. py_tbparse-0.1.1a0/tests/test_datasources.py +63 -0
  23. py_tbparse-0.1.1a0/tests/test_diff.py +68 -0
  24. py_tbparse-0.1.1a0/tests/test_fields.py +50 -0
  25. py_tbparse-0.1.1a0/tests/test_graph.py +70 -0
  26. py_tbparse-0.1.1a0/tests/test_gui_browser.py +228 -0
  27. py_tbparse-0.1.1a0/tests/test_joins.py +74 -0
  28. py_tbparse-0.1.1a0/tests/test_parser.py +81 -0
  29. py_tbparse-0.1.1a0/tests/test_published.py +52 -0
  30. py_tbparse-0.1.1a0/tests/test_relationships.py +62 -0
  31. py_tbparse-0.1.1a0/tests/test_sql.py +97 -0
  32. py_tbparse-0.1.1a0/tests/test_validators.py +62 -0
  33. py_tbparse-0.1.1a0/tests/test_webgui.py +228 -0
  34. py_tbparse-0.1.1a0/tests/test_xml.py +54 -0
  35. py_tbparse-0.1.1a0/twbparser_py/__init__.py +54 -0
  36. py_tbparse-0.1.1a0/twbparser_py/_clean.py +87 -0
  37. py_tbparse-0.1.1a0/twbparser_py/_tables.py +97 -0
  38. py_tbparse-0.1.1a0/twbparser_py/_xml.py +172 -0
  39. py_tbparse-0.1.1a0/twbparser_py/batch.py +62 -0
  40. py_tbparse-0.1.1a0/twbparser_py/calculated_fields.py +113 -0
  41. py_tbparse-0.1.1a0/twbparser_py/cli.py +207 -0
  42. py_tbparse-0.1.1a0/twbparser_py/dashboards.py +101 -0
  43. py_tbparse-0.1.1a0/twbparser_py/datasources.py +257 -0
  44. py_tbparse-0.1.1a0/twbparser_py/diff.py +86 -0
  45. py_tbparse-0.1.1a0/twbparser_py/fields.py +148 -0
  46. py_tbparse-0.1.1a0/twbparser_py/graph.py +81 -0
  47. py_tbparse-0.1.1a0/twbparser_py/joins.py +99 -0
  48. py_tbparse-0.1.1a0/twbparser_py/parser.py +234 -0
  49. py_tbparse-0.1.1a0/twbparser_py/published.py +52 -0
  50. py_tbparse-0.1.1a0/twbparser_py/py.typed +0 -0
  51. py_tbparse-0.1.1a0/twbparser_py/relationships.py +175 -0
  52. py_tbparse-0.1.1a0/twbparser_py/sql.py +76 -0
  53. py_tbparse-0.1.1a0/twbparser_py/validators.py +99 -0
  54. py_tbparse-0.1.1a0/twbparser_py/webgui.py +398 -0
@@ -0,0 +1,231 @@
1
+ # AGENTS.md
2
+
3
+ ## What this is
4
+
5
+ `twbparser_py` is a native Python port of the R package
6
+ [`twbparser`](https://github.com/PrigasG/twbparser) (mirrored at
7
+ `DDSNA/twbparser`): it parses Tableau `.twb`/`.twbx` workbook files into
8
+ `pandas` DataFrames. Pure `lxml` XML parsing — no R runtime, no `rpy2`.
9
+
10
+ Current scope is a **v1 subset** of the original R package's ~50 exported
11
+ functions. Ported: workbook loading (`.twb`/`.twbx`), datasources,
12
+ parameters, raw/calculated fields, joins, relationships (legacy `<relation
13
+ type="join">` and 2020.2+ `<relationships>`), inferred relationships,
14
+ dashboards/dashboard-sheets, `validate_relationships`, custom/initial SQL,
15
+ and published-source detection. **Not yet ported** (v2): formatting,
16
+ tooltips, colors, axes, sorts, dashboard layout/actions, analytics
17
+ helpers (calc complexity, field usage, replication brief), and the
18
+ Shiny-inspector equivalent.
19
+
20
+ Also added, with no R equivalent (Python-native extras -- see "Non-R
21
+ modules" below): Graphviz DOT export of the relationship graph
22
+ (replaces the R package's igraph/ggraph-based plotting with a
23
+ dependency-free alternative), workbook-to-workbook diff, folder/batch
24
+ analysis across many workbooks, and Jupyter rich display.
25
+
26
+ ## Setup
27
+
28
+ No system-wide `pip`; use the project venv.
29
+
30
+ ```bash
31
+ python3 -m venv .venv
32
+ .venv/bin/pip install -e ".[test]"
33
+ ```
34
+
35
+ ## Build / test / lint commands
36
+
37
+ ```bash
38
+ .venv/bin/python -m pytest -q # run the full test suite
39
+ .venv/bin/python -m pytest -q tests/test_joins.py # single file
40
+ .venv/bin/python -m py_compile twbparser_py/*.py # syntax check
41
+ ```
42
+
43
+ Browser (GUI) tests need a one-time setup; without it they skip and the
44
+ rest of the suite still runs:
45
+
46
+ ```bash
47
+ .venv/bin/pip install -e ".[browser]"
48
+ .venv/bin/playwright install chromium
49
+ ./scripts/setup-browser-libs.sh # only on a machine without root
50
+ .venv/bin/python -m pytest -q tests/test_gui_browser.py
51
+ ```
52
+
53
+ There is no linter/formatter configured yet — match existing style
54
+ (no trailing comments unless explaining non-obvious behavior, type hints
55
+ via `from __future__ import annotations`).
56
+
57
+ ## Release / packaging
58
+
59
+ ```bash
60
+ .venv/bin/pip install -e ".[dev]" # build + twine
61
+ rm -rf dist build twbparser_py.egg-info
62
+ .venv/bin/python -m build # produces dist/*.whl and dist/*.tar.gz
63
+ .venv/bin/twine check dist/* # validates metadata/README rendering
64
+ ```
65
+
66
+ Version is single-sourced from `pyproject.toml`'s `[project].version`;
67
+ `twbparser_py.__version__` reads it back via `importlib.metadata` at
68
+ runtime (see `__init__.py`), so don't hardcode a second copy.
69
+
70
+ Before bumping the version for a release: bump `version` in
71
+ `pyproject.toml`, then rebuild and smoke-test the wheel in a
72
+ throwaway venv (`pip install dist/*.whl`, run `twbparser --help` and
73
+ `twbparser-gui --help`, run pytest against an extracted sdist) — this
74
+ catches packaging bugs (missing files, wrong entry points) that an
75
+ editable install won't.
76
+
77
+ CI (`.github/workflows/ci.yml`) runs pytest across Python 3.9–3.13, runs
78
+ the browser GUI tests in real Chromium (`gui` job), and builds+checks the
79
+ distribution on every push/PR. `build` depends on both `test` and `gui`
80
+ — don't drop `gui` from `build`'s `needs`, or a broken-page build can
81
+ succeed and publish an artifact while the one job that would have caught
82
+ it fails off to the side. Release
83
+ (`.github/workflows/release.yml`) publishes to PyPI via trusted
84
+ publishing (OIDC, no stored token) when a GitHub Release is published —
85
+ this needs to be configured once on the PyPI project's "Trusted
86
+ Publishers" settings page before the first release, and needs a
87
+ `release` GitHub Environment created in repo settings.
88
+
89
+ ## Architecture
90
+
91
+ Each `twbparser_py/*.py` module is a direct port of one R source file in
92
+ the upstream package, function-for-function:
93
+
94
+ | Python module | Ported from (R) | Notes |
95
+ |---|---|---|
96
+ | `_clean.py` | `R/utils.R` (`.twb_clean_table`, `.twb_clean_field`, `attr_safe_get`, `.strip_brackets`) | Regex-based name cleaners shared by everything else |
97
+ | `_xml.py` | `R/utils.R` (`twbx_list`, `extract_twb_from_twbx`, `twbx_extract_files`) | `.twbx` zip handling |
98
+ | `datasources.py` | `R/datasources.R` | `extract_named_connections`, `extract_datasource_details`, `extract_parameters` |
99
+ | `fields.py` | `R/fields.R` | `extract_columns_with_table_source`, `infer_implicit_relationships` |
100
+ | `calculated_fields.py` | `R/calculated_fields.R` | `extract_calculated_fields`, `extract_raw_fields` |
101
+ | `joins.py` | `R/joins.R` | `extract_joins` (legacy `<relation type="join">`) |
102
+ | `relationships.py` | `R/relationships.R` | `extract_relations`, `build_object_table_mapping`, `extract_relationships` (2020.2+ model) |
103
+ | `dashboards.py` | `R/dashboard_details.R` (subset) | `list_dashboards`, `dashboard_sheets` |
104
+ | `validators.py` | `R/validators.R` | `validate_relationships` |
105
+ | `sql.py` | `R/sql.R` | `extract_custom_sql` (`twb_custom_sql`), `extract_initial_sql` (`twb_initial_sql`) |
106
+ | `published.py` | `R/published.R` | `extract_published_refs` (`twb_published_refs`) |
107
+ | `parser.py` | `R/twb_parser.R` (`TwbParser` R6 class) | The `TwbParser` façade tying every module together |
108
+
109
+ `parser.py`'s `TwbParser.__init__` mirrors the R6 constructor: it eagerly
110
+ computes every cached DataFrame, guarding each with `_safe_call` (the port
111
+ of R's `safe_call`/`tryCatch`) so a malformed workbook degrades to empty,
112
+ correctly-columned DataFrames instead of raising.
113
+
114
+ ### Non-R modules
115
+
116
+ These have no upstream R source to track — they're either the
117
+ Python-native user-facing layer on top of `TwbParser`, or features added
118
+ beyond the R package's scope:
119
+
120
+ | Python module | What it is |
121
+ |---|---|
122
+ | `_tables.py` | Name → `TwbParser`-accessor registry shared by `cli.py`, `webgui.py`, `diff.py`, and `batch.py`. Adding a new extractor to `TwbParser`? Add it here too so it's automatically available everywhere else. |
123
+ | `cli.py` | `twbparser` command-line entry point, plus the `diff`/`batch` subcommands (dispatched on `sys.argv[1]` before the normal single-workbook argparse parser runs) |
124
+ | `webgui.py` | `twbparser-gui`: stdlib-only (`http.server` + vanilla JS) local browser GUI, no GUI toolkit dependency. Loosely fills the role of the R package's `run_twbparser_app`/Shiny inspector. |
125
+ | `graph.py` | `to_dot()`: Graphviz DOT export of joins/relationships (+ optional inferred, as dashed edges). Replaces the R package's igraph/ggraph-based `plot_dependency_graph`/`plot_relationship_graph` with a dependency-free text format any Graphviz-compatible tool can render. |
126
+ | `diff.py` | `diff_tables()`/`diff_workbooks()`: row-level added/removed diff between two workbooks' same-named table, via `_tables.TABLE_SPECS`. No "changed" classification without a natural key — a changed row shows as one removed + one added row. |
127
+ | `batch.py` | `scan_folder()`: runs one table across every `.twb`/`.twbx` in a directory, concatenated with a `workbook` column. Skips (with a warning) any file that fails to load/extract rather than aborting the batch. |
128
+
129
+ ## Porting conventions (read before adding/modifying a function)
130
+
131
+ 1. **Go to the R source first.** Fetch the corresponding file from
132
+ `github.com/DDSNA/twbparser` (or `PrigasG/twbparser`, same content) via
133
+ the GitHub API/raw URLs — don't guess at Tableau XML schema from
134
+ memory. Port the XPath expressions and fallback logic as literally as
135
+ possible; deviations from the R behavior are bugs, not improvements.
136
+ 2. **XPath parity**: `lxml`'s XPath 1.0 engine supports the same
137
+ functions R's `xml2`/libxml2 use (`local-name()`, `contains()`,
138
+ `count()`), so most R XPath strings can be copied verbatim into
139
+ `.xpath(...)` calls.
140
+ 3. **Empty-input contract**: every extractor returns an empty DataFrame
141
+ with the *correct columns* (not just `pd.DataFrame()`) when there's no
142
+ matching XML — callers (esp. `parser.py`) rely on this. Each module
143
+ exposes its column list as a `_XXX_COLUMNS` constant (e.g.
144
+ `joins._JOIN_COLUMNS`, `datasources._DATASOURCE_COLUMNS`) precisely so
145
+ `parser.py`'s `_safe_call` fallbacks can reuse it instead of retyping
146
+ the list a second time (which drifted out of sync once already).
147
+ 4. **Cleaning helpers**: always route table/field name cleanup through
148
+ `_clean.clean_table` / `_clean.clean_field` / `_clean.strip_brackets`
149
+ rather than re-deriving regexes inline. Same for "is this value
150
+ missing" checks — use `_clean.is_missing(x)`, not a hand-rolled
151
+ `pd.isna()`/`isinstance(x, float)` check (its edge cases, like NaT or
152
+ array-likes, are easy to get subtly wrong per call site).
153
+ 5. **No R dependency, ever.** If a future port needs something R gets
154
+ from `dplyr`/`igraph` for free (e.g. graph layout for
155
+ `plot_dependency_graph`), find a pure-Python equivalent or scope it
156
+ out — don't reach for `rpy2`/subprocess-to-R.
157
+ 6. **XPath string literals**: never build one with
158
+ `f"...='{value}'"` or a `.replace("'", "")` — user/workbook-controlled
159
+ values (a dashboard name, say) can contain `'`, `"`, or both, and
160
+ XPath 1.0 has no in-literal escape character. Use
161
+ `dashboards._xpath_string_literal(value)` (or extend it if a new
162
+ module needs the same thing), which picks a quoting style or falls
163
+ back to `concat()` as needed — verified against real `lxml` evaluation
164
+ for the tricky cases (leading/trailing/consecutive quotes, both quote
165
+ types at once).
166
+
167
+ ## Testing conventions
168
+
169
+ - Fixtures live in `tests/fixtures/` (`test_for_wenjie.twb`,
170
+ `test_for_zip.twbx`) — pulled directly from the R package's
171
+ `inst/extdata/`. Don't regenerate/hand-edit them; if a new fixture is
172
+ needed, pull it from upstream the same way, or build a minimal
173
+ synthetic XML snippet (see `tests/test_joins.py`,
174
+ `tests/test_dashboards.py` for the pattern — many are lifted from the R
175
+ functions' own `@examples` roxygen blocks).
176
+ - `tests/conftest.py` provides `wenjie_xml`, `wenjie_path`,
177
+ `zip_twbx_path` fixtures and an `xml_from_string()` helper. Tests import
178
+ it with `from conftest import xml_from_string` (no `tests/__init__.py`,
179
+ so it's a plain top-level import, not a relative one).
180
+ - Every new extractor function needs: one test against a real fixture (if
181
+ it exercises fixture data) and/or one synthetic-XML test for edge
182
+ cases the fixtures don't cover (empty input, fallback paths).
183
+ - When in doubt about expected output, that's a signal to install R +
184
+ the `twbparser` package and diff against the real function's output on
185
+ the same fixture, rather than guessing.
186
+
187
+ ### Testing the GUI
188
+
189
+ `tests/test_webgui.py` drives the HTTP endpoints directly — it never
190
+ executes the page's JavaScript. That blind spot shipped a real bug: a
191
+ `SyntaxError` in the inline `<script>` (a `\n` in a non-raw Python
192
+ string became a literal newline, splitting a JS string literal across
193
+ two lines) killed the entire script, so no handlers bound and the UI was
194
+ inert — with the whole suite green.
195
+
196
+ So: **any change to `webgui.py`'s `_PAGE` needs a browser test**, in
197
+ `tests/test_gui_browser.py`, which runs the page in real headless
198
+ Chromium via Playwright and fails on any uncaught JS error. The cheap
199
+ structural guards in `test_webgui.py` (unterminated string literals,
200
+ bracket balance) are a backstop, not a substitute.
201
+
202
+ `_PAGE` is declared as `r"""..."""` (a **raw** string) specifically so
203
+ this can't recur: without `r`, any backslash escape meant for the
204
+ *browser* (`\n`, `\t`, a future `\'`) would need doubling in the Python
205
+ source, and forgetting to double it silently corrupts the embedded JS
206
+ instead of erroring. Keep it raw — write JS escapes the normal JS way
207
+ (`'\n'`, not `'\\n'`).
208
+
209
+ Any server-side value spliced into `_PAGE` (currently `TABLE_NAMES` and
210
+ the preloaded workbook path) must go through `_json_for_script()`, not
211
+ bare `json.dumps()`. `json.dumps` doesn't escape `/`, so a value
212
+ containing the literal text `</script>` closes the script tag early in
213
+ the browser's HTML parser — this is real, not theoretical, since the
214
+ workbook path is user-controlled input; see
215
+ `test_preload_path_cannot_break_out_of_script_tag` for a reproduction.
216
+
217
+ `scripts/setup-browser-libs.sh` unpacks Chromium's system libraries into
218
+ a gitignored `.browser-libs/` instead of apt-installing them as root;
219
+ the test fixture picks that directory up automatically via
220
+ `LD_LIBRARY_PATH`. It verifies each download against the checksum
221
+ `apt-get --print-uris` reports for it (SHA256 if offered, MD5Sum as the
222
+ realistic fallback — this environment's apt only emits MD5Sum) before
223
+ extracting; don't remove that step to "simplify" the script. Don't add
224
+ `pytest-playwright` — it's a pytest plugin that imports playwright at
225
+ startup, which makes collection fail for anyone who doesn't have it
226
+ installed.
227
+
228
+ ## Commit / PR conventions
229
+
230
+ Nothing project-specific beyond the harness defaults — see repo commit
231
+ history for message style.
@@ -0,0 +1,25 @@
1
+ MIT License
2
+
3
+ This project is a Python port of logic originally implemented in the R
4
+ package "twbparser" (https://github.com/PrigasG/twbparser),
5
+ Copyright (c) 2025 George Arthur.
6
+
7
+ Copyright (c) 2026 the twbparser-py contributors
8
+
9
+ Permission is hereby granted, free of charge, to any person obtaining a copy
10
+ of this software and associated documentation files (the "Software"), to deal
11
+ in the Software without restriction, including without limitation the rights
12
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
13
+ copies of the Software, and to permit persons to whom the Software is
14
+ furnished to do so, subject to the following conditions:
15
+
16
+ The above copyright notice and this permission notice shall be included in all
17
+ copies or substantial portions of the Software.
18
+
19
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
20
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
21
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
22
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
23
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
24
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
25
+ SOFTWARE.
@@ -0,0 +1,6 @@
1
+ include LICENSE
2
+ include README.md
3
+ include AGENTS.md
4
+ recursive-include tests *.py
5
+ include tests/fixtures/*.twb
6
+ include tests/fixtures/*.twbx
@@ -0,0 +1,168 @@
1
+ Metadata-Version: 2.4
2
+ Name: py-tbparse
3
+ Version: 0.1.1a0
4
+ Summary: Native Python port of the twbparser R package: parse Tableau .twb/.twbx workbooks into pandas DataFrames.
5
+ Author: DDSNA
6
+ License: MIT
7
+ Keywords: tableau,twb,twbx,workbook,parser,pandas
8
+ Classifier: Development Status :: 4 - Beta
9
+ Classifier: Intended Audience :: Developers
10
+ Classifier: License :: OSI Approved :: MIT License
11
+ Classifier: Operating System :: OS Independent
12
+ Classifier: Programming Language :: Python :: 3
13
+ Classifier: Programming Language :: Python :: 3.9
14
+ Classifier: Programming Language :: Python :: 3.10
15
+ Classifier: Programming Language :: Python :: 3.11
16
+ Classifier: Programming Language :: Python :: 3.12
17
+ Classifier: Programming Language :: Python :: 3.13
18
+ Classifier: Topic :: Software Development :: Libraries :: Python Modules
19
+ Classifier: Topic :: Office/Business
20
+ Requires-Python: >=3.9
21
+ Description-Content-Type: text/markdown
22
+ License-File: LICENSE
23
+ Requires-Dist: lxml>=4.9
24
+ Requires-Dist: pandas>=1.5
25
+ Provides-Extra: test
26
+ Requires-Dist: pytest>=7; extra == "test"
27
+ Provides-Extra: browser
28
+ Requires-Dist: playwright>=1.40; extra == "browser"
29
+ Provides-Extra: dev
30
+ Requires-Dist: build>=1.0; extra == "dev"
31
+ Requires-Dist: twine>=5.0; extra == "dev"
32
+ Dynamic: license-file
33
+
34
+ # twbparser-py
35
+
36
+ A native Python port of the [`twbparser`](https://github.com/PrigasG/twbparser)
37
+ R package: parses Tableau `.twb`/`.twbx` workbook files into `pandas`
38
+ DataFrames. No R runtime required — pure `lxml` XML parsing.
39
+
40
+ This is a v1 subset covering the parser's core: workbook loading,
41
+ datasources, parameters, fields, calculated fields, joins, relationships
42
+ (legacy and 2020.2+), inferred relationships, dashboards, relationship
43
+ validation, custom/initial SQL, and published-source detection.
44
+ Formatting/tooltips/colors/axes/sorts, dashboard layout/actions,
45
+ analytics helpers (calc complexity, field usage, replication brief), and
46
+ the Shiny-inspector equivalent are not yet ported.
47
+
48
+ Beyond the R original, this port also adds a few Python-native extras: a
49
+ Graphviz DOT export of the relationship graph, a workbook-to-workbook
50
+ diff, folder/batch analysis across many workbooks, and Jupyter rich
51
+ display (`_repr_html_`).
52
+
53
+ ## Install
54
+
55
+ ```bash
56
+ pip install -e .
57
+ ```
58
+
59
+ ## Usage
60
+
61
+ ```python
62
+ from twbparser_py import TwbParser
63
+
64
+ p = TwbParser("workbook.twb") # or .twbx
65
+ p.get_datasources()
66
+ p.get_fields()
67
+ p.get_calculated_fields()
68
+ p.get_joins()
69
+ p.get_relationships()
70
+ p.get_inferred_relationships()
71
+ p.get_dashboards()
72
+ p.get_dashboard_sheets()
73
+ p.get_custom_sql()
74
+ p.get_initial_sql()
75
+ p.get_published_refs()
76
+ p.get_relationship_graph_dot() # Graphviz DOT string
77
+ p.validate()
78
+ p.get_overview()
79
+ p # in Jupyter: renders get_overview() via _repr_html_
80
+
81
+ from twbparser_py import diff_workbooks, scan_folder
82
+
83
+ diff_workbooks(TwbParser("v1.twb"), TwbParser("v2.twb"), table="datasources")
84
+ scan_folder("./workbooks", table="datasources") # one row per workbook x datasource
85
+ ```
86
+
87
+ ### CLI
88
+
89
+ ```bash
90
+ twbparser workbook.twb # overview (default table)
91
+ twbparser workbook.twb tables # list available tables
92
+ twbparser workbook.twb calculated-fields # print a table
93
+ twbparser workbook.twb fields --format csv -o fields.csv
94
+ twbparser workbook.twbx dashboard-sheets --dashboard "Sales Overview"
95
+ twbparser workbook.twb validate # exit code 2 if invalid
96
+ twbparser workbook.twb graph --include-inferred > relationships.dot
97
+ twbparser diff old.twb new.twb datasources # row-level added/removed
98
+ twbparser batch ./workbooks datasources # one table, every workbook in a folder
99
+ ```
100
+
101
+ Tables: `overview`, `datasources`, `parameters`, `fields`, `raw-fields`,
102
+ `calculated-fields`, `joins`, `relations`, `relationships`,
103
+ `inferred-relationships`, `dashboards`, `dashboard-sheets`,
104
+ `custom-sql`, `initial-sql`, `published-refs`. `--format` is `table`
105
+ (default), `csv`, or `json` (`graph` always prints Graphviz DOT text
106
+ regardless of `--format`). `diff`/`batch` accept most of the same table
107
+ names, minus `graph`/`validate`/`tables`.
108
+
109
+ ### GUI
110
+
111
+ A local, browser-based GUI — standard library only (`http.server` +
112
+ vanilla JS), no GUI toolkit or extra dependency required:
113
+
114
+ ```bash
115
+ twbparser-gui workbook.twb # opens your default browser
116
+ twbparser-gui # opens with an empty path field; paste one and click Load
117
+ twbparser-gui --no-browser --port 8765 # just run the server, e.g. for a headless box
118
+ ```
119
+
120
+ Pick a table from the dropdown, filter `dashboard-sheets` by dashboard,
121
+ toggle "include Parameters" for `calculated-fields`, pick `graph` to
122
+ preview/export a Graphviz DOT digraph of the relationships, and export
123
+ any tabular view as CSV. All state lives server-side in memory for the
124
+ life of the process — it's a single-user local tool, not something to
125
+ expose on a shared network.
126
+
127
+ ## Testing
128
+
129
+ ```bash
130
+ pip install -e ".[test]"
131
+ pytest
132
+ ```
133
+
134
+ Fixtures in `tests/fixtures/` are the same tiny sample workbooks used by
135
+ the original R package's test suite (`inst/extdata/`).
136
+
137
+ The GUI additionally has end-to-end tests that drive the page in a real
138
+ headless Chromium (via Playwright) and fail on any uncaught JavaScript
139
+ error. They're opt-in — without the browser installed they skip and the
140
+ rest of the suite runs normally:
141
+
142
+ ```bash
143
+ pip install -e ".[test,browser]"
144
+ playwright install chromium # add --with-deps if you have root
145
+ ./scripts/setup-browser-libs.sh # no-root alternative to --with-deps
146
+ pytest tests/test_gui_browser.py
147
+ ```
148
+
149
+ ## Background
150
+
151
+ `.twb` is plain XML, and `.twbx` is just a zip wrapper around one, which
152
+ is why parsing it natively in Python — no R, no reverse-engineering —
153
+ was tractable at this scale. Power BI's equivalent format, `.pbix`, is a
154
+ binary container built around the proprietary VertiPaq storage engine,
155
+ which is why reading it programmatically needed dedicated
156
+ reverse-engineering projects like
157
+ [PBIXRay](https://github.com/Hugoberry/pbixray) and
158
+ [pbi-tools](https://github.com/pbi-tools/pbi-tools). The comparison
159
+ isn't one-sided, though: Tableau's own official Python tooling for
160
+ *server* automation
161
+ ([`tableauserverclient`](https://pypi.org/project/tableauserverclient/),
162
+ `tabcmd`) is more mature and more open than anything Microsoft ships for
163
+ Power BI's REST API.
164
+
165
+ ## Credit
166
+
167
+ Ported from the R implementation by George Arthur
168
+ ([`PrigasG/twbparser`](https://github.com/PrigasG/twbparser)), MIT licensed.
@@ -0,0 +1,135 @@
1
+ # twbparser-py
2
+
3
+ A native Python port of the [`twbparser`](https://github.com/PrigasG/twbparser)
4
+ R package: parses Tableau `.twb`/`.twbx` workbook files into `pandas`
5
+ DataFrames. No R runtime required — pure `lxml` XML parsing.
6
+
7
+ This is a v1 subset covering the parser's core: workbook loading,
8
+ datasources, parameters, fields, calculated fields, joins, relationships
9
+ (legacy and 2020.2+), inferred relationships, dashboards, relationship
10
+ validation, custom/initial SQL, and published-source detection.
11
+ Formatting/tooltips/colors/axes/sorts, dashboard layout/actions,
12
+ analytics helpers (calc complexity, field usage, replication brief), and
13
+ the Shiny-inspector equivalent are not yet ported.
14
+
15
+ Beyond the R original, this port also adds a few Python-native extras: a
16
+ Graphviz DOT export of the relationship graph, a workbook-to-workbook
17
+ diff, folder/batch analysis across many workbooks, and Jupyter rich
18
+ display (`_repr_html_`).
19
+
20
+ ## Install
21
+
22
+ ```bash
23
+ pip install -e .
24
+ ```
25
+
26
+ ## Usage
27
+
28
+ ```python
29
+ from twbparser_py import TwbParser
30
+
31
+ p = TwbParser("workbook.twb") # or .twbx
32
+ p.get_datasources()
33
+ p.get_fields()
34
+ p.get_calculated_fields()
35
+ p.get_joins()
36
+ p.get_relationships()
37
+ p.get_inferred_relationships()
38
+ p.get_dashboards()
39
+ p.get_dashboard_sheets()
40
+ p.get_custom_sql()
41
+ p.get_initial_sql()
42
+ p.get_published_refs()
43
+ p.get_relationship_graph_dot() # Graphviz DOT string
44
+ p.validate()
45
+ p.get_overview()
46
+ p # in Jupyter: renders get_overview() via _repr_html_
47
+
48
+ from twbparser_py import diff_workbooks, scan_folder
49
+
50
+ diff_workbooks(TwbParser("v1.twb"), TwbParser("v2.twb"), table="datasources")
51
+ scan_folder("./workbooks", table="datasources") # one row per workbook x datasource
52
+ ```
53
+
54
+ ### CLI
55
+
56
+ ```bash
57
+ twbparser workbook.twb # overview (default table)
58
+ twbparser workbook.twb tables # list available tables
59
+ twbparser workbook.twb calculated-fields # print a table
60
+ twbparser workbook.twb fields --format csv -o fields.csv
61
+ twbparser workbook.twbx dashboard-sheets --dashboard "Sales Overview"
62
+ twbparser workbook.twb validate # exit code 2 if invalid
63
+ twbparser workbook.twb graph --include-inferred > relationships.dot
64
+ twbparser diff old.twb new.twb datasources # row-level added/removed
65
+ twbparser batch ./workbooks datasources # one table, every workbook in a folder
66
+ ```
67
+
68
+ Tables: `overview`, `datasources`, `parameters`, `fields`, `raw-fields`,
69
+ `calculated-fields`, `joins`, `relations`, `relationships`,
70
+ `inferred-relationships`, `dashboards`, `dashboard-sheets`,
71
+ `custom-sql`, `initial-sql`, `published-refs`. `--format` is `table`
72
+ (default), `csv`, or `json` (`graph` always prints Graphviz DOT text
73
+ regardless of `--format`). `diff`/`batch` accept most of the same table
74
+ names, minus `graph`/`validate`/`tables`.
75
+
76
+ ### GUI
77
+
78
+ A local, browser-based GUI — standard library only (`http.server` +
79
+ vanilla JS), no GUI toolkit or extra dependency required:
80
+
81
+ ```bash
82
+ twbparser-gui workbook.twb # opens your default browser
83
+ twbparser-gui # opens with an empty path field; paste one and click Load
84
+ twbparser-gui --no-browser --port 8765 # just run the server, e.g. for a headless box
85
+ ```
86
+
87
+ Pick a table from the dropdown, filter `dashboard-sheets` by dashboard,
88
+ toggle "include Parameters" for `calculated-fields`, pick `graph` to
89
+ preview/export a Graphviz DOT digraph of the relationships, and export
90
+ any tabular view as CSV. All state lives server-side in memory for the
91
+ life of the process — it's a single-user local tool, not something to
92
+ expose on a shared network.
93
+
94
+ ## Testing
95
+
96
+ ```bash
97
+ pip install -e ".[test]"
98
+ pytest
99
+ ```
100
+
101
+ Fixtures in `tests/fixtures/` are the same tiny sample workbooks used by
102
+ the original R package's test suite (`inst/extdata/`).
103
+
104
+ The GUI additionally has end-to-end tests that drive the page in a real
105
+ headless Chromium (via Playwright) and fail on any uncaught JavaScript
106
+ error. They're opt-in — without the browser installed they skip and the
107
+ rest of the suite runs normally:
108
+
109
+ ```bash
110
+ pip install -e ".[test,browser]"
111
+ playwright install chromium # add --with-deps if you have root
112
+ ./scripts/setup-browser-libs.sh # no-root alternative to --with-deps
113
+ pytest tests/test_gui_browser.py
114
+ ```
115
+
116
+ ## Background
117
+
118
+ `.twb` is plain XML, and `.twbx` is just a zip wrapper around one, which
119
+ is why parsing it natively in Python — no R, no reverse-engineering —
120
+ was tractable at this scale. Power BI's equivalent format, `.pbix`, is a
121
+ binary container built around the proprietary VertiPaq storage engine,
122
+ which is why reading it programmatically needed dedicated
123
+ reverse-engineering projects like
124
+ [PBIXRay](https://github.com/Hugoberry/pbixray) and
125
+ [pbi-tools](https://github.com/pbi-tools/pbi-tools). The comparison
126
+ isn't one-sided, though: Tableau's own official Python tooling for
127
+ *server* automation
128
+ ([`tableauserverclient`](https://pypi.org/project/tableauserverclient/),
129
+ `tabcmd`) is more mature and more open than anything Microsoft ships for
130
+ Power BI's REST API.
131
+
132
+ ## Credit
133
+
134
+ Ported from the R implementation by George Arthur
135
+ ([`PrigasG/twbparser`](https://github.com/PrigasG/twbparser)), MIT licensed.