exstruct 0.2.80__tar.gz → 0.3.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (31) hide show
  1. {exstruct-0.2.80 → exstruct-0.3.0}/PKG-INFO +22 -12
  2. {exstruct-0.2.80 → exstruct-0.3.0}/README.md +21 -11
  3. {exstruct-0.2.80 → exstruct-0.3.0}/pyproject.toml +128 -119
  4. {exstruct-0.2.80 → exstruct-0.3.0}/src/exstruct/__init__.py +23 -12
  5. {exstruct-0.2.80 → exstruct-0.3.0}/src/exstruct/cli/main.py +20 -0
  6. exstruct-0.3.0/src/exstruct/core/backends/__init__.py +7 -0
  7. exstruct-0.3.0/src/exstruct/core/backends/base.py +38 -0
  8. exstruct-0.3.0/src/exstruct/core/backends/com_backend.py +226 -0
  9. exstruct-0.3.0/src/exstruct/core/backends/openpyxl_backend.py +179 -0
  10. {exstruct-0.2.80 → exstruct-0.3.0}/src/exstruct/core/cells.py +964 -483
  11. exstruct-0.3.0/src/exstruct/core/integrate.py +52 -0
  12. exstruct-0.3.0/src/exstruct/core/logging_utils.py +16 -0
  13. exstruct-0.3.0/src/exstruct/core/modeling.py +74 -0
  14. exstruct-0.3.0/src/exstruct/core/pipeline.py +696 -0
  15. exstruct-0.3.0/src/exstruct/core/ranges.py +48 -0
  16. exstruct-0.3.0/src/exstruct/core/workbook.py +114 -0
  17. {exstruct-0.2.80 → exstruct-0.3.0}/src/exstruct/engine.py +42 -123
  18. {exstruct-0.2.80 → exstruct-0.3.0}/src/exstruct/errors.py +46 -35
  19. {exstruct-0.2.80 → exstruct-0.3.0}/src/exstruct/io/__init__.py +72 -132
  20. exstruct-0.3.0/src/exstruct/io/serialize.py +112 -0
  21. {exstruct-0.2.80 → exstruct-0.3.0}/src/exstruct/models/__init__.py +11 -3
  22. {exstruct-0.2.80 → exstruct-0.3.0}/src/exstruct/render/__init__.py +3 -7
  23. exstruct-0.2.80/src/exstruct/core/integrate.py +0 -388
  24. {exstruct-0.2.80 → exstruct-0.3.0}/LICENSE +0 -0
  25. {exstruct-0.2.80 → exstruct-0.3.0}/src/exstruct/cli/availability.py +0 -0
  26. {exstruct-0.2.80 → exstruct-0.3.0}/src/exstruct/core/__init__.py +0 -0
  27. {exstruct-0.2.80 → exstruct-0.3.0}/src/exstruct/core/charts.py +0 -0
  28. {exstruct-0.2.80 → exstruct-0.3.0}/src/exstruct/core/shapes.py +0 -0
  29. {exstruct-0.2.80 → exstruct-0.3.0}/src/exstruct/models/maps.py +0 -0
  30. {exstruct-0.2.80 → exstruct-0.3.0}/src/exstruct/models/types.py +0 -0
  31. {exstruct-0.2.80 → exstruct-0.3.0}/src/exstruct/py.typed +0 -0
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.3
2
2
  Name: exstruct
3
- Version: 0.2.80
3
+ Version: 0.3.0
4
4
  Summary: Excel to structured JSON (tables, shapes, charts) for LLM/RAG pipelines
5
5
  Keywords: excel,structure,data,exstruct
6
6
  Author: harumiWeb
@@ -61,12 +61,12 @@ Description-Content-Type: text/markdown
61
61
 
62
62
  ExStruct reads Excel workbooks and outputs structured data (cells, table candidates, shapes, charts, print areas/views, auto page-break areas, hyperlinks) as JSON by default, with optional YAML/TOON formats. It targets both COM/Excel environments (rich extraction) and non-COM environments (cells + table candidates + print areas), with tunable detection heuristics and multiple output modes to fit LLM/RAG pipelines.
63
63
 
64
- [日本版README](README.ja.md)
64
+ [日本版 README](README.ja.md)
65
65
 
66
66
  ## Features
67
67
 
68
68
  - **Excel → Structured JSON**: cells, shapes, charts, table candidates, print areas/views, and auto page-break areas per sheet.
69
- - **Output modes**: `light` (cells + table candidates + print areas; no COM, shapes/charts empty), `standard` (texted shapes + arrows, charts, print areas), `verbose` (all shapes with width/height, charts with size, print areas). Verbose also emits cell hyperlinks. Size output is flag-controlled.
69
+ - **Output modes**: `light` (cells + table candidates + print areas; no COM, shapes/charts empty), `standard` (texted shapes + arrows, charts, print areas), `verbose` (all shapes with width/height, charts with size, print areas). Verbose also emits cell hyperlinks and `colors_map`. Size output is flag-controlled.
70
70
  - **Auto page-break export (COM only)**: capture Excel-computed auto page breaks and write per-area JSON/YAML/TOON when requested (CLI option appears only when COM is available).
71
71
  - **Formats**: JSON (compact by default, `--pretty` available), YAML, TOON (optional dependencies).
72
72
  - **Table detection tuning**: adjust heuristics at runtime via API.
@@ -189,7 +189,7 @@ Use higher thresholds to reduce false positives; lower them if true tables are m
189
189
 
190
190
  - **light**: cells + table candidates (no COM needed).
191
191
  - **standard**: texted shapes + arrows, charts (COM if available), table candidates. Hyperlinks are off unless `include_cell_links=True`.
192
- - **verbose**: all shapes (with width/height), charts, table candidates, and cell hyperlinks.
192
+ - **verbose**: all shapes (with width/height), charts, table candidates, cell hyperlinks, and `colors_map`.
193
193
 
194
194
  ## Error Handling / Fallbacks
195
195
 
@@ -218,7 +218,6 @@ To show how well exstruct can structure Excel, we parse a workbook that combines
218
218
  (Screenshot below is the actual sample Excel sheet)
219
219
  ![Sample Excel](/docs/assets/demo_sheet.png)
220
220
  Sample workbook: `sample/sample.xlsx`
221
- Sample workbook: `sample/sample.xlsx`
222
221
 
223
222
  ### 1. Input: Excel Sheet Overview
224
223
 
@@ -274,12 +273,14 @@ Below is a **shortened JSON output example** from parsing this Excel workbook.
274
273
  ],
275
274
  "shapes": [
276
275
  {
276
+ "id": 1,
277
277
  "text": "開始",
278
278
  "l": 148,
279
279
  "t": 220,
280
280
  "type": "AutoShape-FlowchartProcess"
281
281
  },
282
282
  {
283
+ "id": 2,
283
284
  "text": "入力データ読み込み",
284
285
  "l": 132,
285
286
  "t": 282,
@@ -291,6 +292,8 @@ Below is a **shortened JSON output example** from parsing this Excel workbook.
291
292
  "type": "AutoShape-Mixed",
292
293
  "begin_arrow_style": 1,
293
294
  "end_arrow_style": 2,
295
+ "begin_id": 1,
296
+ "end_id": 2,
294
297
  "direction": "N"
295
298
  },
296
299
  ...
@@ -374,14 +377,14 @@ flowchart TD
374
377
 
375
378
  A --> B
376
379
  B --> C
377
- C -- no --> D
378
- C -- yes --> E
380
+ C -->|yes| D
381
+ C --> H
382
+ D --> E
379
383
  E --> F
380
- F -- yes --> E
381
- F -- no --> G
382
- G --> H
383
- H -- yes --> I
384
- H -- no --> J
384
+ F --> G
385
+ G -->|yes| I
386
+ G -->|no| J
387
+ H --> J
385
388
  I --> J
386
389
  ```
387
390
  ````
@@ -390,6 +393,12 @@ From this we can see:
390
393
 
391
394
  **exstruct's JSON is already in a format that AI can read and reason over directly.**
392
395
 
396
+ Other LLM inference samples using this library can be found in the following directory:
397
+
398
+ - [Basic Excel](sample/basic/)
399
+ - [Flowchart](sample/flowchart/)
400
+ - [Gantt Chart](sample/gantt_chart/)
401
+
393
402
  ### 4. Summary
394
403
 
395
404
  This benchmark confirms exstruct can:
@@ -414,6 +423,7 @@ ExStruct is used primarily as a **library**, not a service.
414
423
  - Forking and internal modification are expected in enterprise use
415
424
 
416
425
  This project is suitable for teams that:
426
+
417
427
  - need transparency over black-box tools
418
428
  - are comfortable maintaining internal forks if necessary
419
429
 
@@ -6,12 +6,12 @@
6
6
 
7
7
  ExStruct reads Excel workbooks and outputs structured data (cells, table candidates, shapes, charts, print areas/views, auto page-break areas, hyperlinks) as JSON by default, with optional YAML/TOON formats. It targets both COM/Excel environments (rich extraction) and non-COM environments (cells + table candidates + print areas), with tunable detection heuristics and multiple output modes to fit LLM/RAG pipelines.
8
8
 
9
- [日本版README](README.ja.md)
9
+ [日本版 README](README.ja.md)
10
10
 
11
11
  ## Features
12
12
 
13
13
  - **Excel → Structured JSON**: cells, shapes, charts, table candidates, print areas/views, and auto page-break areas per sheet.
14
- - **Output modes**: `light` (cells + table candidates + print areas; no COM, shapes/charts empty), `standard` (texted shapes + arrows, charts, print areas), `verbose` (all shapes with width/height, charts with size, print areas). Verbose also emits cell hyperlinks. Size output is flag-controlled.
14
+ - **Output modes**: `light` (cells + table candidates + print areas; no COM, shapes/charts empty), `standard` (texted shapes + arrows, charts, print areas), `verbose` (all shapes with width/height, charts with size, print areas). Verbose also emits cell hyperlinks and `colors_map`. Size output is flag-controlled.
15
15
  - **Auto page-break export (COM only)**: capture Excel-computed auto page breaks and write per-area JSON/YAML/TOON when requested (CLI option appears only when COM is available).
16
16
  - **Formats**: JSON (compact by default, `--pretty` available), YAML, TOON (optional dependencies).
17
17
  - **Table detection tuning**: adjust heuristics at runtime via API.
@@ -134,7 +134,7 @@ Use higher thresholds to reduce false positives; lower them if true tables are m
134
134
 
135
135
  - **light**: cells + table candidates (no COM needed).
136
136
  - **standard**: texted shapes + arrows, charts (COM if available), table candidates. Hyperlinks are off unless `include_cell_links=True`.
137
- - **verbose**: all shapes (with width/height), charts, table candidates, and cell hyperlinks.
137
+ - **verbose**: all shapes (with width/height), charts, table candidates, cell hyperlinks, and `colors_map`.
138
138
 
139
139
  ## Error Handling / Fallbacks
140
140
 
@@ -163,7 +163,6 @@ To show how well exstruct can structure Excel, we parse a workbook that combines
163
163
  (Screenshot below is the actual sample Excel sheet)
164
164
  ![Sample Excel](/docs/assets/demo_sheet.png)
165
165
  Sample workbook: `sample/sample.xlsx`
166
- Sample workbook: `sample/sample.xlsx`
167
166
 
168
167
  ### 1. Input: Excel Sheet Overview
169
168
 
@@ -219,12 +218,14 @@ Below is a **shortened JSON output example** from parsing this Excel workbook.
219
218
  ],
220
219
  "shapes": [
221
220
  {
221
+ "id": 1,
222
222
  "text": "開始",
223
223
  "l": 148,
224
224
  "t": 220,
225
225
  "type": "AutoShape-FlowchartProcess"
226
226
  },
227
227
  {
228
+ "id": 2,
228
229
  "text": "入力データ読み込み",
229
230
  "l": 132,
230
231
  "t": 282,
@@ -236,6 +237,8 @@ Below is a **shortened JSON output example** from parsing this Excel workbook.
236
237
  "type": "AutoShape-Mixed",
237
238
  "begin_arrow_style": 1,
238
239
  "end_arrow_style": 2,
240
+ "begin_id": 1,
241
+ "end_id": 2,
239
242
  "direction": "N"
240
243
  },
241
244
  ...
@@ -319,14 +322,14 @@ flowchart TD
319
322
 
320
323
  A --> B
321
324
  B --> C
322
- C -- no --> D
323
- C -- yes --> E
325
+ C -->|yes| D
326
+ C --> H
327
+ D --> E
324
328
  E --> F
325
- F -- yes --> E
326
- F -- no --> G
327
- G --> H
328
- H -- yes --> I
329
- H -- no --> J
329
+ F --> G
330
+ G -->|yes| I
331
+ G -->|no| J
332
+ H --> J
330
333
  I --> J
331
334
  ```
332
335
  ````
@@ -335,6 +338,12 @@ From this we can see:
335
338
 
336
339
  **exstruct's JSON is already in a format that AI can read and reason over directly.**
337
340
 
341
+ Other LLM inference samples using this library can be found in the following directory:
342
+
343
+ - [Basic Excel](sample/basic/)
344
+ - [Flowchart](sample/flowchart/)
345
+ - [Gantt Chart](sample/gantt_chart/)
346
+
338
347
  ### 4. Summary
339
348
 
340
349
  This benchmark confirms exstruct can:
@@ -359,6 +368,7 @@ ExStruct is used primarily as a **library**, not a service.
359
368
  - Forking and internal modification are expected in enterprise use
360
369
 
361
370
  This project is suitable for teams that:
371
+
362
372
  - need transparency over black-box tools
363
373
  - are comfortable maintaining internal forks if necessary
364
374
 
@@ -1,119 +1,128 @@
1
- [project]
2
- name = "exstruct"
3
- version = "0.2.80"
4
- description = "Excel to structured JSON (tables, shapes, charts) for LLM/RAG pipelines"
5
- readme = "README.md"
6
- license = { file = "LICENSE" }
7
- keywords = ["excel", "structure", "data", "exstruct"]
8
- authors = [
9
- { name = "harumiWeb"}
10
- ]
11
- requires-python = ">=3.11"
12
- dependencies = [
13
- "numpy>=2.3.5",
14
- "openpyxl>=3.1.5",
15
- "pandas>=2.3.3",
16
- "pydantic>=2.12.5",
17
- "scipy>=1.16.3",
18
- "xlwings>=0.33.16",
19
- ]
20
-
21
- [build-system]
22
- requires = ["uv_build>=0.8.4,<0.9.0"]
23
- build-backend = "uv_build"
24
-
25
- [dependency-groups]
26
- dev = [
27
- "mkdocs-material>=9.7.0",
28
- "mkdocstrings-python>=2.0.1",
29
- "mypy>=1.19.0",
30
- "pre-commit>=4.5.0",
31
- "pytest>=9.0.1",
32
- "pytest-cov>=7.0.0",
33
- "pytest-mock>=3.15.1",
34
- "ruff>=0.14.8",
35
- ]
36
-
37
- [project.optional-dependencies]
38
- yaml = ["pyyaml>=6.0.3"]
39
- toon = ["python-toon>=0.1.3"]
40
- render = ["pypdfium2>=5.1.0", "Pillow>=12.0.0"]
41
-
42
- [project.scripts]
43
- exstruct = "exstruct.cli.main:main"
44
-
45
- [project.urls]
46
- Homepage = "https://harumiweb.github.io/exstruct/"
47
- Repository = "https://github.com/harumiWeb/exstruct"
48
- Issues = "https://github.com/harumiWeb/exstruct/issues"
49
- Documentation = "https://harumiweb.github.io/exstruct/"
50
-
51
- [tool.coverage.run]
52
- omit = [
53
- "tests/*",
54
- "*/test_*.py",
55
- "*/gen_py/*",
56
- ]
57
-
58
- [tool.ruff]
59
- target-version = "py311"
60
- src = ["exstruct"]
61
-
62
- select = [
63
- "E", # pycodestyle errors
64
- "W", # pycodestyle warnings
65
- "F", # pyflakes
66
- "I", # import sorting
67
- "UP", # pyupgrade
68
- "B", # flake8-bugbear
69
- "N", # naming
70
- "C90", # complexity
71
- "A", # flake8-builtins
72
- "ANN", # type annotations
73
- ]
74
-
75
- ignore = [
76
- "E501", # 行長は許容(Excel JSON は長くなりがち)
77
- "B008", # Pydantic の default_factory を誤検知するため
78
- "ANN101", # self に型を要求されてしまうため
79
- "ANN102", # cls も同様
80
- ]
81
-
82
- fix = true
83
-
84
- # 型ヒントのスタイル
85
- [tool.ruff.lint]
86
- extend-select = ["ANN"]
87
-
88
- # import の並び替え設定
89
- [tool.ruff.isort]
90
- combine-as-imports = true
91
- known-first-party = ["exstruct"]
92
- force-sort-within-sections = true
93
-
94
- # 複雑度チェック(関数の最大複雑度)
95
- [tool.ruff.mccabe]
96
- max-complexity = 12
97
-
98
- [tool.ruff.per-file-ignores]
99
- "tests/**/*.py" = ["N802", "N803", "N806"]
100
-
101
-
102
- [tool.mypy]
103
- packages = ["exstruct"]
104
- python_version = "3.11"
105
-
106
- # 外部ライブラリは一切チェックしない
107
- ignore_missing_imports = true
108
-
109
- # 自作コードは厳密にチェックする
110
- strict = true
111
-
112
- # Pydantic v2 向け
113
- plugins = ["pydantic.mypy"]
114
-
115
- [tool.pytest.ini_options]
116
- markers = [
117
- "com: requires Excel COM (Windows + Excel)",
118
- "render: requires Excel COM and pypdfium2; set RUN_RENDER_SMOKE=1 to enable",
119
- ]
1
+ [project]
2
+ name = "exstruct"
3
+ version = "0.3.0"
4
+ description = "Excel to structured JSON (tables, shapes, charts) for LLM/RAG pipelines"
5
+ readme = "README.md"
6
+ license = { file = "LICENSE" }
7
+ keywords = ["excel", "structure", "data", "exstruct"]
8
+ authors = [
9
+ { name = "harumiWeb"}
10
+ ]
11
+ requires-python = ">=3.11"
12
+ dependencies = [
13
+ "numpy>=2.3.5",
14
+ "openpyxl>=3.1.5",
15
+ "pandas>=2.3.3",
16
+ "pydantic>=2.12.5",
17
+ "scipy>=1.16.3",
18
+ "xlwings>=0.33.16",
19
+ ]
20
+
21
+ [build-system]
22
+ requires = ["uv_build>=0.8.4,<0.9.0"]
23
+ build-backend = "uv_build"
24
+
25
+ [dependency-groups]
26
+ dev = [
27
+ "mkdocs-material>=9.7.0",
28
+ "mkdocstrings-python>=2.0.1",
29
+ "mypy>=1.19.0",
30
+ "pre-commit>=4.5.0",
31
+ "pytest>=9.0.1",
32
+ "pytest-cov>=7.0.0",
33
+ "pytest-mock>=3.15.1",
34
+ "ruff>=0.14.8",
35
+ "taskipy>=1.14.1",
36
+ ]
37
+
38
+ [project.optional-dependencies]
39
+ yaml = ["pyyaml>=6.0.3"]
40
+ toon = ["python-toon>=0.1.3"]
41
+ render = ["pypdfium2>=5.1.0", "Pillow>=12.0.0"]
42
+
43
+ [project.scripts]
44
+ exstruct = "exstruct.cli.main:main"
45
+
46
+ [project.urls]
47
+ Homepage = "https://harumiweb.github.io/exstruct/"
48
+ Repository = "https://github.com/harumiWeb/exstruct"
49
+ Issues = "https://github.com/harumiWeb/exstruct/issues"
50
+ Documentation = "https://harumiweb.github.io/exstruct/"
51
+
52
+ [tool.coverage.run]
53
+ omit = [
54
+ "tests/*",
55
+ "*/test_*.py",
56
+ "*/gen_py/*",
57
+ ]
58
+
59
+ [tool.ruff]
60
+ target-version = "py311"
61
+ src = ["exstruct"]
62
+
63
+ select = [
64
+ "E", # pycodestyle errors
65
+ "W", # pycodestyle warnings
66
+ "F", # pyflakes
67
+ "I", # import sorting
68
+ "UP", # pyupgrade
69
+ "B", # flake8-bugbear
70
+ "N", # naming
71
+ "C90", # complexity
72
+ "A", # flake8-builtins
73
+ "ANN", # type annotations
74
+ ]
75
+
76
+ ignore = [
77
+ "E501", # 行長は許容(Excel JSON は長くなりがち)
78
+ "B008", # Pydantic の default_factory を誤検知するため
79
+ "ANN101", # self に型を要求されてしまうため
80
+ "ANN102", # cls も同様
81
+ ]
82
+
83
+ fix = true
84
+
85
+ # 型ヒントのスタイル
86
+ [tool.ruff.lint]
87
+ extend-select = ["ANN"]
88
+
89
+ # import の並び替え設定
90
+ [tool.ruff.isort]
91
+ combine-as-imports = true
92
+ known-first-party = ["exstruct"]
93
+ force-sort-within-sections = true
94
+
95
+ # 複雑度チェック(関数の最大複雑度)
96
+ [tool.ruff.mccabe]
97
+ max-complexity = 12
98
+
99
+ [tool.ruff.per-file-ignores]
100
+ "tests/**/*.py" = ["N802", "N803", "N806"]
101
+
102
+
103
+ [tool.mypy]
104
+ packages = ["exstruct"]
105
+ python_version = "3.11"
106
+
107
+ # 外部ライブラリは一切チェックしない
108
+ ignore_missing_imports = true
109
+
110
+ # 自作コードは厳密にチェックする
111
+ strict = true
112
+
113
+ # Pydantic v2 向け
114
+ plugins = ["pydantic.mypy"]
115
+
116
+ [tool.pytest.ini_options]
117
+ markers = [
118
+ "com: requires Excel COM (Windows + Excel)",
119
+ "render: requires Excel COM and pypdfium2; set RUN_RENDER_SMOKE=1 to enable",
120
+ ]
121
+
122
+ [tool.taskipy.tasks]
123
+ ruff = "ruff check ."
124
+ ruff-fix = "ruff check . --fix"
125
+ mypy = "mypy src/exstruct --strict"
126
+ test = "pytest -vv --cov=exstruct --cov-report=term-missing --cov-report=xml"
127
+ docs = "mkdocs serve"
128
+ build-docs = "mkdocs build && python scripts/gen_json_schema.py && python scripts/gen_model_docs.py"
@@ -7,9 +7,11 @@ from typing import Literal, TextIO
7
7
  from .core.cells import set_table_detection_params
8
8
  from .core.integrate import extract_workbook
9
9
  from .engine import (
10
+ ColorsOptions,
10
11
  DestinationOptions,
11
12
  ExStructEngine,
12
13
  FilterOptions,
14
+ FormatOptions,
13
15
  OutputOptions,
14
16
  StructOptions,
15
17
  )
@@ -75,7 +77,9 @@ __all__ = [
75
77
  "StructOptions",
76
78
  "OutputOptions",
77
79
  "FilterOptions",
80
+ "FormatOptions",
78
81
  "DestinationOptions",
82
+ "ColorsOptions",
79
83
  "serialize_workbook",
80
84
  "export_auto_page_breaks",
81
85
  ]
@@ -93,7 +97,7 @@ def extract(file_path: str | Path, mode: ExtractionMode = "standard") -> Workboo
93
97
  mode: "light" / "standard" / "verbose"
94
98
  - light: cells + table detection only (no COM, shapes/charts empty). Print areas via openpyxl.
95
99
  - standard: texted shapes + arrows + charts (COM if available), print areas included. Shape/chart size is kept but hidden by default in output.
96
- - verbose: all shapes (including textless) with size, charts with size.
100
+ - verbose: all shapes (including textless) with size, charts with size, and colors_map.
97
101
 
98
102
  Returns:
99
103
  WorkbookData containing sheets, rows, shapes, charts, and print areas.
@@ -110,8 +114,13 @@ def extract(file_path: str | Path, mode: ExtractionMode = "standard") -> Workboo
110
114
  ['A1:B5']
111
115
  """
112
116
  include_links = True if mode == "verbose" else False
117
+ include_colors_map = True if mode == "verbose" else None
113
118
  engine = ExStructEngine(
114
- options=StructOptions(mode=mode, include_cell_links=include_links)
119
+ options=StructOptions(
120
+ mode=mode,
121
+ include_cell_links=include_links,
122
+ include_colors_map=include_colors_map,
123
+ )
115
124
  )
116
125
  return engine.extract(file_path, mode=mode)
117
126
 
@@ -358,16 +367,18 @@ def process_excel(
358
367
  engine = ExStructEngine(
359
368
  options=StructOptions(mode=mode),
360
369
  output=OutputOptions(
361
- fmt=out_fmt,
362
- pretty=pretty,
363
- indent=indent,
364
- sheets_dir=sheets_dir,
365
- print_areas_dir=print_areas_dir,
366
- auto_page_breaks_dir=auto_page_breaks_dir,
367
- include_print_areas=None if mode == "light" else True,
368
- include_shape_size=True if mode == "verbose" else False,
369
- include_chart_size=True if mode == "verbose" else False,
370
- stream=stream,
370
+ format=FormatOptions(fmt=out_fmt, pretty=pretty, indent=indent),
371
+ filters=FilterOptions(
372
+ include_print_areas=None if mode == "light" else True,
373
+ include_shape_size=True if mode == "verbose" else False,
374
+ include_chart_size=True if mode == "verbose" else False,
375
+ ),
376
+ destinations=DestinationOptions(
377
+ sheets_dir=sheets_dir,
378
+ print_areas_dir=print_areas_dir,
379
+ auto_page_breaks_dir=auto_page_breaks_dir,
380
+ stream=stream,
381
+ ),
371
382
  ),
372
383
  )
373
384
  engine.process(
@@ -2,11 +2,30 @@ from __future__ import annotations
2
2
 
3
3
  import argparse
4
4
  from pathlib import Path
5
+ import sys
5
6
 
6
7
  from exstruct import process_excel
7
8
  from exstruct.cli.availability import ComAvailability, get_com_availability
8
9
 
9
10
 
11
+ def _ensure_utf8_stdout() -> None:
12
+ """Reconfigure stdout to UTF-8 when supported.
13
+
14
+ Windows consoles default to cp932 and can raise encoding errors when piping
15
+ non-ASCII characters. Reconfiguring prevents failures without affecting
16
+ environments that already default to UTF-8.
17
+ """
18
+
19
+ stdout = sys.stdout
20
+ if not hasattr(stdout, "reconfigure"):
21
+ return
22
+ reconfigure = stdout.reconfigure
23
+ try:
24
+ reconfigure(encoding="utf-8", errors="replace")
25
+ except (AttributeError, ValueError):
26
+ return
27
+
28
+
10
29
  def _add_auto_page_breaks_argument(
11
30
  parser: argparse.ArgumentParser, availability: ComAvailability
12
31
  ) -> None:
@@ -102,6 +121,7 @@ def main(argv: list[str] | None = None) -> int:
102
121
  Returns:
103
122
  Exit code (0 for success, 1 for failure).
104
123
  """
124
+ _ensure_utf8_stdout()
105
125
  parser = build_parser()
106
126
  args = parser.parse_args(argv)
107
127
 
@@ -0,0 +1,7 @@
1
+ from __future__ import annotations
2
+
3
+ from .base import Backend
4
+ from .com_backend import ComBackend
5
+ from .openpyxl_backend import OpenpyxlBackend
6
+
7
+ __all__ = ["Backend", "ComBackend", "OpenpyxlBackend"]
@@ -0,0 +1,38 @@
1
+ from __future__ import annotations
2
+
3
+ from dataclasses import dataclass
4
+ from typing import Protocol
5
+
6
+ from ...models import CellRow, PrintArea
7
+ from ..cells import WorkbookColorsMap
8
+
9
+ CellData = dict[str, list[CellRow]]
10
+ PrintAreaData = dict[str, list[PrintArea]]
11
+
12
+
13
+ @dataclass(frozen=True)
14
+ class BackendConfig:
15
+ """Configuration options shared across backends.
16
+
17
+ Attributes:
18
+ include_default_background: Whether to include default background colors.
19
+ ignore_colors: Optional set of color keys to ignore.
20
+ """
21
+
22
+ include_default_background: bool
23
+ ignore_colors: set[str] | None
24
+
25
+
26
+ class Backend(Protocol):
27
+ """Protocol for backend implementations."""
28
+
29
+ def extract_cells(self, *, include_links: bool) -> CellData:
30
+ """Extract cell rows from the workbook."""
31
+
32
+ def extract_print_areas(self) -> PrintAreaData:
33
+ """Extract print areas from the workbook."""
34
+
35
+ def extract_colors_map(
36
+ self, *, include_default_background: bool, ignore_colors: set[str] | None
37
+ ) -> WorkbookColorsMap | None:
38
+ """Extract colors map from the workbook."""