docxsurgeon 0.1.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- docxsurgeon-0.1.0/LICENSE +21 -0
- docxsurgeon-0.1.0/PKG-INFO +203 -0
- docxsurgeon-0.1.0/README.md +177 -0
- docxsurgeon-0.1.0/pyproject.toml +41 -0
- docxsurgeon-0.1.0/setup.cfg +4 -0
- docxsurgeon-0.1.0/src/docxsurgeon/__init__.py +4 -0
- docxsurgeon-0.1.0/src/docxsurgeon/package.py +228 -0
- docxsurgeon-0.1.0/src/docxsurgeon.egg-info/PKG-INFO +203 -0
- docxsurgeon-0.1.0/src/docxsurgeon.egg-info/SOURCES.txt +13 -0
- docxsurgeon-0.1.0/src/docxsurgeon.egg-info/dependency_links.txt +1 -0
- docxsurgeon-0.1.0/src/docxsurgeon.egg-info/requires.txt +1 -0
- docxsurgeon-0.1.0/src/docxsurgeon.egg-info/top_level.txt +1 -0
- docxsurgeon-0.1.0/tests/test_chart_and_symbol.py +89 -0
- docxsurgeon-0.1.0/tests/test_embedded.py +75 -0
- docxsurgeon-0.1.0/tests/test_package.py +83 -0
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
MIT License
|
|
2
|
+
|
|
3
|
+
Copyright (c) 2026 txxcat
|
|
4
|
+
|
|
5
|
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
|
6
|
+
of this software and associated documentation files (the "Software"), to deal
|
|
7
|
+
in the Software without restriction, including without limitation the rights
|
|
8
|
+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
|
9
|
+
copies of the Software, and to permit persons to whom the Software is
|
|
10
|
+
furnished to do so, subject to the following conditions:
|
|
11
|
+
|
|
12
|
+
The above copyright notice and this permission notice shall be included in all
|
|
13
|
+
copies or substantial portions of the Software.
|
|
14
|
+
|
|
15
|
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
|
16
|
+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
|
17
|
+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
|
18
|
+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
|
19
|
+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
|
20
|
+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
|
21
|
+
SOFTWARE.
|
|
@@ -0,0 +1,203 @@
|
|
|
1
|
+
Metadata-Version: 2.4
|
|
2
|
+
Name: docxsurgeon
|
|
3
|
+
Version: 0.1.0
|
|
4
|
+
Summary: 对 docx 做 zip 包级手术:替换/编辑嵌入的 xlsx、更新图表数据范围、修复复选框打印符号
|
|
5
|
+
Author: txxcat
|
|
6
|
+
License: MIT
|
|
7
|
+
Project-URL: Homepage, https://github.com/txxcat/docxsurgeon
|
|
8
|
+
Keywords: docx,ooxml,embedded,excel,chart,word,surgery
|
|
9
|
+
Classifier: Development Status :: 3 - Alpha
|
|
10
|
+
Classifier: Intended Audience :: Developers
|
|
11
|
+
Classifier: License :: OSI Approved :: MIT License
|
|
12
|
+
Classifier: Operating System :: OS Independent
|
|
13
|
+
Classifier: Programming Language :: Python :: 3
|
|
14
|
+
Classifier: Programming Language :: Python :: 3.9
|
|
15
|
+
Classifier: Programming Language :: Python :: 3.10
|
|
16
|
+
Classifier: Programming Language :: Python :: 3.11
|
|
17
|
+
Classifier: Programming Language :: Python :: 3.12
|
|
18
|
+
Classifier: Programming Language :: Python :: 3.13
|
|
19
|
+
Classifier: Topic :: Office/Business :: Office Suites
|
|
20
|
+
Classifier: Topic :: Software Development :: Libraries :: Python Modules
|
|
21
|
+
Requires-Python: >=3.9
|
|
22
|
+
Description-Content-Type: text/markdown
|
|
23
|
+
License-File: LICENSE
|
|
24
|
+
Requires-Dist: openpyxl>=3.0
|
|
25
|
+
Dynamic: license-file
|
|
26
|
+
|
|
27
|
+
# docxsurgeon
|
|
28
|
+
|
|
29
|
+
[](https://github.com/txxcat/docxsurgeon/actions/workflows/ci.yml)
|
|
30
|
+
[](LICENSE)
|
|
31
|
+
[](https://www.python.org)
|
|
32
|
+
[](#roadmap)
|
|
33
|
+
|
|
34
|
+
**Operate on the *inside* of .docx files — replace embedded Excel, update chart ranges, fix checkbox glyphs. All in memory.**
|
|
35
|
+
|
|
36
|
+
**对 docx 做 zip 包级"手术":替换/编辑嵌入的 xlsx、更新图表数据引用范围、修复复选框打印符号。全程内存操作,不落地临时目录。**
|
|
37
|
+
|
|
38
|
+
## 为什么需要它
|
|
39
|
+
|
|
40
|
+
OOXML 文件(docx/xlsx/pptx)本质是 zip 包,里面藏着 `word/embeddings/*.xlsx`(嵌入的 Excel)、`word/charts/chart*.xml`(图表定义)、`xl/tables/*.xml`(Excel 表格范围)。这些**包内结构**在主流库里无人认领:
|
|
41
|
+
|
|
42
|
+
| 需求 | python-docx | docxtpl | openpyxl | docxsurgeon |
|
|
43
|
+
|---|---|---|---|---|
|
|
44
|
+
| 替换 docx 内嵌入的 xlsx | ❌ | ⚠️ 需走 Jinja2 render 流程 | ❌(只认独立文件) | ✅ |
|
|
45
|
+
| **编辑**嵌入 xlsx 的单元格 | ❌ | ❌ | ❌ | ✅ 内存直改,退出自动写回 |
|
|
46
|
+
| 改嵌入 xlsx 里 Excel 表格范围 | ❌ | ❌ | ❌ | ✅ |
|
|
47
|
+
| 更新图表 XML 数据引用行数 | ❌ | ❌ | ❌ | ✅ |
|
|
48
|
+
| Unicode ☑ → Wingdings 符号 run(打印修复) | ❌ | ❌ | ❌ | ✅ |
|
|
49
|
+
| 文档模板变量填充 | ❌ | ✅ | — | ❌(不抢 docxtpl 的活) |
|
|
50
|
+
| 文档模型读写(段落/表格) | ✅ | ✅ | — | ❌(不抢 python-docx 的活) |
|
|
51
|
+
|
|
52
|
+
一句话定位:**python-docx 管文档模型,docxtpl 管模板填充,docxsurgeon 管"已有文档的包级修改"。**
|
|
53
|
+
|
|
54
|
+
核心逻辑提取自一套在广电播控机房生产环境运行多年的报表工具,经过真实 Word / WPS 文档的长期验证。
|
|
55
|
+
|
|
56
|
+
## 安装
|
|
57
|
+
|
|
58
|
+
要求 **Python >= 3.9**。
|
|
59
|
+
|
|
60
|
+
**从 PyPI(待发布)**
|
|
61
|
+
|
|
62
|
+
> 本项目尚未在 PyPI 正式发布,0.1.0 上线后会启用下面的命令。当前请使用下方的源码安装方式。
|
|
63
|
+
|
|
64
|
+
```bash
|
|
65
|
+
pip install docxsurgeon
|
|
66
|
+
```
|
|
67
|
+
|
|
68
|
+
**从源码安装(当前阶段推荐)**
|
|
69
|
+
|
|
70
|
+
```bash
|
|
71
|
+
pip install "git+https://github.com/txxcat/docxsurgeon.git"
|
|
72
|
+
```
|
|
73
|
+
|
|
74
|
+
**开发模式**
|
|
75
|
+
|
|
76
|
+
```bash
|
|
77
|
+
git clone https://github.com/txxcat/docxsurgeon
|
|
78
|
+
cd docxsurgeon
|
|
79
|
+
pip install -e .
|
|
80
|
+
```
|
|
81
|
+
|
|
82
|
+
依赖说明:运行时仅依赖 `openpyxl>=3.0`(编辑嵌入 xlsx 时用到,已随安装自动拉取)。其余功能(部件读写、图表范围改写、符号修复)不依赖任何第三方库。
|
|
83
|
+
|
|
84
|
+
## 快速上手
|
|
85
|
+
|
|
86
|
+
```python
|
|
87
|
+
from docxsurgeon import DocxPackage, set_table_range
|
|
88
|
+
|
|
89
|
+
pkg = DocxPackage("周报模板.docx")
|
|
90
|
+
|
|
91
|
+
# 1) 直接编辑包内嵌入的 Excel(内存中,无临时目录)
|
|
92
|
+
with pkg.open_embedded_xlsx("word/embeddings/Microsoft_Excel____1.xlsx") as wb:
|
|
93
|
+
ws = wb.active
|
|
94
|
+
ws["A2"] = "9月"
|
|
95
|
+
ws["B2"] = 12
|
|
96
|
+
set_table_range(ws, "A1:C8") # 同步 Excel 表格范围
|
|
97
|
+
|
|
98
|
+
# 2) 图表引用行数随数据条数伸缩(Sheet1!$B$2:$B$20 -> $B$2:$B$8)
|
|
99
|
+
pkg.set_chart_data_rows("word/charts/chart2.xml", last_row=8)
|
|
100
|
+
|
|
101
|
+
# 3) 替换嵌入文件本体
|
|
102
|
+
pkg.replace_embedded("word/embeddings/Microsoft_Word_Document3.docx", "新附件.docx")
|
|
103
|
+
|
|
104
|
+
# 4) 修复 ☑ 打印乱码:替换为 Wingdings 2 符号 run
|
|
105
|
+
pkg.fix_checkbox_glyph()
|
|
106
|
+
|
|
107
|
+
pkg.save("输出.docx") # 未修改的部件字节级原样保留
|
|
108
|
+
```
|
|
109
|
+
|
|
110
|
+
## 进阶示例
|
|
111
|
+
|
|
112
|
+
**探索文档结构** —— 动刀前先看看包里有什么:
|
|
113
|
+
|
|
114
|
+
```python
|
|
115
|
+
pkg = DocxPackage("周报模板.docx")
|
|
116
|
+
print(pkg.list_parts()) # 全部部件,按原始 zip 顺序
|
|
117
|
+
print(pkg.list_embeddings()) # 只列 word/embeddings/ 下的嵌入文件
|
|
118
|
+
print(pkg.has_part("word/document.xml")) # True
|
|
119
|
+
```
|
|
120
|
+
|
|
121
|
+
**部件内文本替换** —— 把某个 XML 部件当纯文本做子串替换,返回替换次数:
|
|
122
|
+
|
|
123
|
+
```python
|
|
124
|
+
n = pkg.replace_text_in_part(
|
|
125
|
+
"word/charts/chart2.xml",
|
|
126
|
+
"Sheet1!$B$2:$B$20",
|
|
127
|
+
"Sheet1!$B$2:$B$8",
|
|
128
|
+
)
|
|
129
|
+
print(f"改了 {n} 处")
|
|
130
|
+
```
|
|
131
|
+
|
|
132
|
+
**失败安全** —— `open_embedded_xlsx` 的 with 块里若抛异常,本次对嵌入簿的修改会被整体丢弃,包保持原样:
|
|
133
|
+
|
|
134
|
+
```python
|
|
135
|
+
try:
|
|
136
|
+
with pkg.open_embedded_xlsx("word/embeddings/data.xlsx") as wb:
|
|
137
|
+
wb.active["A1"] = "新值"
|
|
138
|
+
raise RuntimeError("模拟中途出错")
|
|
139
|
+
except RuntimeError:
|
|
140
|
+
pass
|
|
141
|
+
# 此时嵌入簿未被破坏,可安全地改做别的处理或直接 save()
|
|
142
|
+
```
|
|
143
|
+
|
|
144
|
+
## API
|
|
145
|
+
|
|
146
|
+
### `DocxPackage(source)`
|
|
147
|
+
|
|
148
|
+
`source` 为文件路径或二进制文件对象。
|
|
149
|
+
|
|
150
|
+
| 方法 | 说明 |
|
|
151
|
+
|---|---|
|
|
152
|
+
| `list_parts()` / `has_part(name)` / `read_part(name)` | zip 部件级读取 |
|
|
153
|
+
| `replace_part(name, data, create=False)` | 整体替换部件(bytes/str) |
|
|
154
|
+
| `replace_text_in_part(name, old, new)` | 部件内文本替换,返回次数 |
|
|
155
|
+
| `list_embeddings()` | 列出 `word/embeddings/` 下的全部嵌入文件 |
|
|
156
|
+
| `replace_embedded(name, source)` | 用本地文件或字节替换嵌入文件 |
|
|
157
|
+
| `open_embedded_xlsx(name)` | **上下文管理器**:yield openpyxl Workbook,正常退出自动写回,异常退出放弃修改 |
|
|
158
|
+
| `set_chart_data_rows(chart_part, last_row, sheet="Sheet1", start_row=2)` | 改写图表数据引用范围,返回改写数 |
|
|
159
|
+
| `fix_checkbox_glyph(part, search="☑", symbol_font="Wingdings 2", symbol_char="0052", ...)` | Unicode 勾选符 → 符号 run |
|
|
160
|
+
| `save(path)` | 写回磁盘 |
|
|
161
|
+
|
|
162
|
+
### 模块函数
|
|
163
|
+
|
|
164
|
+
- `set_table_range(ws, ref)` —— 在 `open_embedded_xlsx` 块内修改 Excel 表格(ListObject)范围。
|
|
165
|
+
- `__version__` —— 当前版本号(与 `package.py` 单一来源一致)。
|
|
166
|
+
|
|
167
|
+
## 设计要点
|
|
168
|
+
|
|
169
|
+
- **全内存**:部件读写走 `BytesIO`,没有临时目录、没有 `os.chdir`、没有竞态清理。
|
|
170
|
+
- **保真**:`save()` 按原 zip 顺序写入;未修改的部件字节级不变(测试保证)。
|
|
171
|
+
- **失败安全**:`open_embedded_xlsx` 中抛异常则放弃本次修改,包保持原样。
|
|
172
|
+
- **惰性依赖**:只有用到 `open_embedded_xlsx` 才需要 openpyxl。
|
|
173
|
+
|
|
174
|
+
## 已知边界
|
|
175
|
+
|
|
176
|
+
- 图表的缓存数据(`c:strCache` / `c:numCache`)不会同步改写——Word 打开时会从嵌入工作簿刷新。若你的场景要求"不刷新也立即显示新值",参考 roadmap。
|
|
177
|
+
- 嵌入的 PDF/OLE 对象(`oleObjectNNN.bin`)不支持替换(OLE 是编码格式)。
|
|
178
|
+
- `open_embedded_xlsx` 依赖 openpyxl 往返,openpyxl 不建模的 xlsx 特性会丢失(与直接用 openpyxl 编辑同一文件的风险一致)。
|
|
179
|
+
|
|
180
|
+
## Roadmap
|
|
181
|
+
|
|
182
|
+
- [ ] 上传 PyPI 正式版(当前未发布,`pip install` 需走源码)
|
|
183
|
+
- [ ] 图表缓存数据(strCache/numCache)双写
|
|
184
|
+
- [ ] 嵌入 xlsx 的图表部件(`word/charts` 引用嵌入簿的完整链路更新)
|
|
185
|
+
- [ ] pptx / xlsx 宿主包的一等支持(当前 API 以 docx 命名,机制通用)
|
|
186
|
+
- [ ] CLI:`python -m docxsurgeon list/extract/replace ...`
|
|
187
|
+
|
|
188
|
+
## 开发 & 贡献
|
|
189
|
+
|
|
190
|
+
```bash
|
|
191
|
+
git clone https://github.com/txxcat/docxsurgeon
|
|
192
|
+
cd docxsurgeon
|
|
193
|
+
pip install -e . pytest
|
|
194
|
+
pytest
|
|
195
|
+
```
|
|
196
|
+
|
|
197
|
+
仓库通过 GitHub Actions 跑 CI(见 `.github/workflows/ci.yml`,每次 push/PR 自动跑测试)。
|
|
198
|
+
|
|
199
|
+
想参与贡献?请先阅读 [CONTRIBUTING.md](CONTRIBUTING.md)。
|
|
200
|
+
|
|
201
|
+
## License
|
|
202
|
+
|
|
203
|
+
[MIT](LICENSE)
|
|
@@ -0,0 +1,177 @@
|
|
|
1
|
+
# docxsurgeon
|
|
2
|
+
|
|
3
|
+
[](https://github.com/txxcat/docxsurgeon/actions/workflows/ci.yml)
|
|
4
|
+
[](LICENSE)
|
|
5
|
+
[](https://www.python.org)
|
|
6
|
+
[](#roadmap)
|
|
7
|
+
|
|
8
|
+
**Operate on the *inside* of .docx files — replace embedded Excel, update chart ranges, fix checkbox glyphs. All in memory.**
|
|
9
|
+
|
|
10
|
+
**对 docx 做 zip 包级"手术":替换/编辑嵌入的 xlsx、更新图表数据引用范围、修复复选框打印符号。全程内存操作,不落地临时目录。**
|
|
11
|
+
|
|
12
|
+
## 为什么需要它
|
|
13
|
+
|
|
14
|
+
OOXML 文件(docx/xlsx/pptx)本质是 zip 包,里面藏着 `word/embeddings/*.xlsx`(嵌入的 Excel)、`word/charts/chart*.xml`(图表定义)、`xl/tables/*.xml`(Excel 表格范围)。这些**包内结构**在主流库里无人认领:
|
|
15
|
+
|
|
16
|
+
| 需求 | python-docx | docxtpl | openpyxl | docxsurgeon |
|
|
17
|
+
|---|---|---|---|---|
|
|
18
|
+
| 替换 docx 内嵌入的 xlsx | ❌ | ⚠️ 需走 Jinja2 render 流程 | ❌(只认独立文件) | ✅ |
|
|
19
|
+
| **编辑**嵌入 xlsx 的单元格 | ❌ | ❌ | ❌ | ✅ 内存直改,退出自动写回 |
|
|
20
|
+
| 改嵌入 xlsx 里 Excel 表格范围 | ❌ | ❌ | ❌ | ✅ |
|
|
21
|
+
| 更新图表 XML 数据引用行数 | ❌ | ❌ | ❌ | ✅ |
|
|
22
|
+
| Unicode ☑ → Wingdings 符号 run(打印修复) | ❌ | ❌ | ❌ | ✅ |
|
|
23
|
+
| 文档模板变量填充 | ❌ | ✅ | — | ❌(不抢 docxtpl 的活) |
|
|
24
|
+
| 文档模型读写(段落/表格) | ✅ | ✅ | — | ❌(不抢 python-docx 的活) |
|
|
25
|
+
|
|
26
|
+
一句话定位:**python-docx 管文档模型,docxtpl 管模板填充,docxsurgeon 管"已有文档的包级修改"。**
|
|
27
|
+
|
|
28
|
+
核心逻辑提取自一套在广电播控机房生产环境运行多年的报表工具,经过真实 Word / WPS 文档的长期验证。
|
|
29
|
+
|
|
30
|
+
## 安装
|
|
31
|
+
|
|
32
|
+
要求 **Python >= 3.9**。
|
|
33
|
+
|
|
34
|
+
**从 PyPI(待发布)**
|
|
35
|
+
|
|
36
|
+
> 本项目尚未在 PyPI 正式发布,0.1.0 上线后会启用下面的命令。当前请使用下方的源码安装方式。
|
|
37
|
+
|
|
38
|
+
```bash
|
|
39
|
+
pip install docxsurgeon
|
|
40
|
+
```
|
|
41
|
+
|
|
42
|
+
**从源码安装(当前阶段推荐)**
|
|
43
|
+
|
|
44
|
+
```bash
|
|
45
|
+
pip install "git+https://github.com/txxcat/docxsurgeon.git"
|
|
46
|
+
```
|
|
47
|
+
|
|
48
|
+
**开发模式**
|
|
49
|
+
|
|
50
|
+
```bash
|
|
51
|
+
git clone https://github.com/txxcat/docxsurgeon
|
|
52
|
+
cd docxsurgeon
|
|
53
|
+
pip install -e .
|
|
54
|
+
```
|
|
55
|
+
|
|
56
|
+
依赖说明:运行时仅依赖 `openpyxl>=3.0`(编辑嵌入 xlsx 时用到,已随安装自动拉取)。其余功能(部件读写、图表范围改写、符号修复)不依赖任何第三方库。
|
|
57
|
+
|
|
58
|
+
## 快速上手
|
|
59
|
+
|
|
60
|
+
```python
|
|
61
|
+
from docxsurgeon import DocxPackage, set_table_range
|
|
62
|
+
|
|
63
|
+
pkg = DocxPackage("周报模板.docx")
|
|
64
|
+
|
|
65
|
+
# 1) 直接编辑包内嵌入的 Excel(内存中,无临时目录)
|
|
66
|
+
with pkg.open_embedded_xlsx("word/embeddings/Microsoft_Excel____1.xlsx") as wb:
|
|
67
|
+
ws = wb.active
|
|
68
|
+
ws["A2"] = "9月"
|
|
69
|
+
ws["B2"] = 12
|
|
70
|
+
set_table_range(ws, "A1:C8") # 同步 Excel 表格范围
|
|
71
|
+
|
|
72
|
+
# 2) 图表引用行数随数据条数伸缩(Sheet1!$B$2:$B$20 -> $B$2:$B$8)
|
|
73
|
+
pkg.set_chart_data_rows("word/charts/chart2.xml", last_row=8)
|
|
74
|
+
|
|
75
|
+
# 3) 替换嵌入文件本体
|
|
76
|
+
pkg.replace_embedded("word/embeddings/Microsoft_Word_Document3.docx", "新附件.docx")
|
|
77
|
+
|
|
78
|
+
# 4) 修复 ☑ 打印乱码:替换为 Wingdings 2 符号 run
|
|
79
|
+
pkg.fix_checkbox_glyph()
|
|
80
|
+
|
|
81
|
+
pkg.save("输出.docx") # 未修改的部件字节级原样保留
|
|
82
|
+
```
|
|
83
|
+
|
|
84
|
+
## 进阶示例
|
|
85
|
+
|
|
86
|
+
**探索文档结构** —— 动刀前先看看包里有什么:
|
|
87
|
+
|
|
88
|
+
```python
|
|
89
|
+
pkg = DocxPackage("周报模板.docx")
|
|
90
|
+
print(pkg.list_parts()) # 全部部件,按原始 zip 顺序
|
|
91
|
+
print(pkg.list_embeddings()) # 只列 word/embeddings/ 下的嵌入文件
|
|
92
|
+
print(pkg.has_part("word/document.xml")) # True
|
|
93
|
+
```
|
|
94
|
+
|
|
95
|
+
**部件内文本替换** —— 把某个 XML 部件当纯文本做子串替换,返回替换次数:
|
|
96
|
+
|
|
97
|
+
```python
|
|
98
|
+
n = pkg.replace_text_in_part(
|
|
99
|
+
"word/charts/chart2.xml",
|
|
100
|
+
"Sheet1!$B$2:$B$20",
|
|
101
|
+
"Sheet1!$B$2:$B$8",
|
|
102
|
+
)
|
|
103
|
+
print(f"改了 {n} 处")
|
|
104
|
+
```
|
|
105
|
+
|
|
106
|
+
**失败安全** —— `open_embedded_xlsx` 的 with 块里若抛异常,本次对嵌入簿的修改会被整体丢弃,包保持原样:
|
|
107
|
+
|
|
108
|
+
```python
|
|
109
|
+
try:
|
|
110
|
+
with pkg.open_embedded_xlsx("word/embeddings/data.xlsx") as wb:
|
|
111
|
+
wb.active["A1"] = "新值"
|
|
112
|
+
raise RuntimeError("模拟中途出错")
|
|
113
|
+
except RuntimeError:
|
|
114
|
+
pass
|
|
115
|
+
# 此时嵌入簿未被破坏,可安全地改做别的处理或直接 save()
|
|
116
|
+
```
|
|
117
|
+
|
|
118
|
+
## API
|
|
119
|
+
|
|
120
|
+
### `DocxPackage(source)`
|
|
121
|
+
|
|
122
|
+
`source` 为文件路径或二进制文件对象。
|
|
123
|
+
|
|
124
|
+
| 方法 | 说明 |
|
|
125
|
+
|---|---|
|
|
126
|
+
| `list_parts()` / `has_part(name)` / `read_part(name)` | zip 部件级读取 |
|
|
127
|
+
| `replace_part(name, data, create=False)` | 整体替换部件(bytes/str) |
|
|
128
|
+
| `replace_text_in_part(name, old, new)` | 部件内文本替换,返回次数 |
|
|
129
|
+
| `list_embeddings()` | 列出 `word/embeddings/` 下的全部嵌入文件 |
|
|
130
|
+
| `replace_embedded(name, source)` | 用本地文件或字节替换嵌入文件 |
|
|
131
|
+
| `open_embedded_xlsx(name)` | **上下文管理器**:yield openpyxl Workbook,正常退出自动写回,异常退出放弃修改 |
|
|
132
|
+
| `set_chart_data_rows(chart_part, last_row, sheet="Sheet1", start_row=2)` | 改写图表数据引用范围,返回改写数 |
|
|
133
|
+
| `fix_checkbox_glyph(part, search="☑", symbol_font="Wingdings 2", symbol_char="0052", ...)` | Unicode 勾选符 → 符号 run |
|
|
134
|
+
| `save(path)` | 写回磁盘 |
|
|
135
|
+
|
|
136
|
+
### 模块函数
|
|
137
|
+
|
|
138
|
+
- `set_table_range(ws, ref)` —— 在 `open_embedded_xlsx` 块内修改 Excel 表格(ListObject)范围。
|
|
139
|
+
- `__version__` —— 当前版本号(与 `package.py` 单一来源一致)。
|
|
140
|
+
|
|
141
|
+
## 设计要点
|
|
142
|
+
|
|
143
|
+
- **全内存**:部件读写走 `BytesIO`,没有临时目录、没有 `os.chdir`、没有竞态清理。
|
|
144
|
+
- **保真**:`save()` 按原 zip 顺序写入;未修改的部件字节级不变(测试保证)。
|
|
145
|
+
- **失败安全**:`open_embedded_xlsx` 中抛异常则放弃本次修改,包保持原样。
|
|
146
|
+
- **惰性依赖**:只有用到 `open_embedded_xlsx` 才需要 openpyxl。
|
|
147
|
+
|
|
148
|
+
## 已知边界
|
|
149
|
+
|
|
150
|
+
- 图表的缓存数据(`c:strCache` / `c:numCache`)不会同步改写——Word 打开时会从嵌入工作簿刷新。若你的场景要求"不刷新也立即显示新值",参考 roadmap。
|
|
151
|
+
- 嵌入的 PDF/OLE 对象(`oleObjectNNN.bin`)不支持替换(OLE 是编码格式)。
|
|
152
|
+
- `open_embedded_xlsx` 依赖 openpyxl 往返,openpyxl 不建模的 xlsx 特性会丢失(与直接用 openpyxl 编辑同一文件的风险一致)。
|
|
153
|
+
|
|
154
|
+
## Roadmap
|
|
155
|
+
|
|
156
|
+
- [ ] 上传 PyPI 正式版(当前未发布,`pip install` 需走源码)
|
|
157
|
+
- [ ] 图表缓存数据(strCache/numCache)双写
|
|
158
|
+
- [ ] 嵌入 xlsx 的图表部件(`word/charts` 引用嵌入簿的完整链路更新)
|
|
159
|
+
- [ ] pptx / xlsx 宿主包的一等支持(当前 API 以 docx 命名,机制通用)
|
|
160
|
+
- [ ] CLI:`python -m docxsurgeon list/extract/replace ...`
|
|
161
|
+
|
|
162
|
+
## 开发 & 贡献
|
|
163
|
+
|
|
164
|
+
```bash
|
|
165
|
+
git clone https://github.com/txxcat/docxsurgeon
|
|
166
|
+
cd docxsurgeon
|
|
167
|
+
pip install -e . pytest
|
|
168
|
+
pytest
|
|
169
|
+
```
|
|
170
|
+
|
|
171
|
+
仓库通过 GitHub Actions 跑 CI(见 `.github/workflows/ci.yml`,每次 push/PR 自动跑测试)。
|
|
172
|
+
|
|
173
|
+
想参与贡献?请先阅读 [CONTRIBUTING.md](CONTRIBUTING.md)。
|
|
174
|
+
|
|
175
|
+
## License
|
|
176
|
+
|
|
177
|
+
[MIT](LICENSE)
|
|
@@ -0,0 +1,41 @@
|
|
|
1
|
+
[build-system]
|
|
2
|
+
requires = ["setuptools>=61"]
|
|
3
|
+
build-backend = "setuptools.build_meta"
|
|
4
|
+
|
|
5
|
+
[project]
|
|
6
|
+
name = "docxsurgeon"
|
|
7
|
+
# 版本号单一来源:只改 src/docxsurgeon/package.py 里的 __version__,这里自动读取
|
|
8
|
+
dynamic = ["version"]
|
|
9
|
+
description = "对 docx 做 zip 包级手术:替换/编辑嵌入的 xlsx、更新图表数据范围、修复复选框打印符号"
|
|
10
|
+
readme = "README.md"
|
|
11
|
+
license = { text = "MIT" }
|
|
12
|
+
authors = [{ name = "txxcat" }]
|
|
13
|
+
requires-python = ">=3.9"
|
|
14
|
+
dependencies = ["openpyxl>=3.0"]
|
|
15
|
+
keywords = ["docx", "ooxml", "embedded", "excel", "chart", "word", "surgery"]
|
|
16
|
+
classifiers = [
|
|
17
|
+
"Development Status :: 3 - Alpha",
|
|
18
|
+
"Intended Audience :: Developers",
|
|
19
|
+
"License :: OSI Approved :: MIT License",
|
|
20
|
+
"Operating System :: OS Independent",
|
|
21
|
+
"Programming Language :: Python :: 3",
|
|
22
|
+
"Programming Language :: Python :: 3.9",
|
|
23
|
+
"Programming Language :: Python :: 3.10",
|
|
24
|
+
"Programming Language :: Python :: 3.11",
|
|
25
|
+
"Programming Language :: Python :: 3.12",
|
|
26
|
+
"Programming Language :: Python :: 3.13",
|
|
27
|
+
"Topic :: Office/Business :: Office Suites",
|
|
28
|
+
"Topic :: Software Development :: Libraries :: Python Modules",
|
|
29
|
+
]
|
|
30
|
+
|
|
31
|
+
[project.urls]
|
|
32
|
+
Homepage = "https://github.com/txxcat/docxsurgeon"
|
|
33
|
+
|
|
34
|
+
[tool.setuptools.packages.find]
|
|
35
|
+
where = ["src"]
|
|
36
|
+
|
|
37
|
+
[tool.setuptools.dynamic]
|
|
38
|
+
version = { attr = "docxsurgeon.package.__version__" }
|
|
39
|
+
|
|
40
|
+
[tool.pytest.ini_options]
|
|
41
|
+
testpaths = ["tests"]
|
|
@@ -0,0 +1,228 @@
|
|
|
1
|
+
"""docxsurgeon:对 OOXML 文件(docx/xlsx)做 zip 包级"手术"。
|
|
2
|
+
|
|
3
|
+
定位:python-docx 管"文档模型",docxtpl 管"模板填充",
|
|
4
|
+
本包管"已有文档的包级修改"——替换/编辑嵌入文件、更新图表
|
|
5
|
+
数据引用范围、修复复选框打印符号。全部在内存中完成,
|
|
6
|
+
不落地临时目录,未修改的部件字节级原样保留。
|
|
7
|
+
|
|
8
|
+
核心逻辑提取自一套在广电播控机房生产环境运行多年的报表工具,
|
|
9
|
+
经过真实 Word/WPS 文档的长期验证。
|
|
10
|
+
"""
|
|
11
|
+
from __future__ import annotations
|
|
12
|
+
|
|
13
|
+
import io
|
|
14
|
+
import re
|
|
15
|
+
import zipfile
|
|
16
|
+
from contextlib import contextmanager
|
|
17
|
+
|
|
18
|
+
__version__ = "0.1.0"
|
|
19
|
+
|
|
20
|
+
EMBEDDINGS_DIR = "word/embeddings/"
|
|
21
|
+
|
|
22
|
+
|
|
23
|
+
class DocxPackage:
|
|
24
|
+
"""把一个 docx(或任意 OOXML zip 包)加载进内存做部件级修改。
|
|
25
|
+
|
|
26
|
+
用法::
|
|
27
|
+
|
|
28
|
+
pkg = DocxPackage("周报模板.docx")
|
|
29
|
+
with pkg.open_embedded_xlsx("word/embeddings/Microsoft_Excel____.xlsx") as wb:
|
|
30
|
+
ws = wb.active
|
|
31
|
+
ws["A2"] = "9月"
|
|
32
|
+
set_table_range(ws, "A1:C8")
|
|
33
|
+
pkg.set_chart_data_rows("word/charts/chart2.xml", last_row=8)
|
|
34
|
+
pkg.fix_checkbox_glyph()
|
|
35
|
+
pkg.save("输出.docx")
|
|
36
|
+
"""
|
|
37
|
+
|
|
38
|
+
def __init__(self, source):
|
|
39
|
+
"""source 可以是文件路径或已打开的二进制文件对象。"""
|
|
40
|
+
self._order = [] # 保持 zip 内部件的原始顺序
|
|
41
|
+
self._parts = {} # arcname -> bytes
|
|
42
|
+
with zipfile.ZipFile(source) as zf:
|
|
43
|
+
for info in zf.infolist():
|
|
44
|
+
if info.is_dir():
|
|
45
|
+
continue
|
|
46
|
+
self._order.append(info.filename)
|
|
47
|
+
self._parts[info.filename] = zf.read(info.filename)
|
|
48
|
+
|
|
49
|
+
# ------------------------------------------------------------------
|
|
50
|
+
# 部件(zip 条目)级操作
|
|
51
|
+
# ------------------------------------------------------------------
|
|
52
|
+
|
|
53
|
+
def list_parts(self):
|
|
54
|
+
"""返回包内全部部件名(按原始 zip 顺序)。"""
|
|
55
|
+
return list(self._order)
|
|
56
|
+
|
|
57
|
+
def has_part(self, name):
|
|
58
|
+
return name in self._parts
|
|
59
|
+
|
|
60
|
+
def read_part(self, name):
|
|
61
|
+
"""按部件名读取原始字节。"""
|
|
62
|
+
self._require_part(name)
|
|
63
|
+
return self._parts[name]
|
|
64
|
+
|
|
65
|
+
def replace_part(self, name, data, create=False):
|
|
66
|
+
"""整体替换一个部件的内容。data 为 bytes 或 str(按 UTF-8 编码)。
|
|
67
|
+
|
|
68
|
+
create=True 时允许写入原包中不存在的新部件(追加到包尾)。
|
|
69
|
+
"""
|
|
70
|
+
if not isinstance(data, (bytes, bytearray)):
|
|
71
|
+
data = str(data).encode("utf-8")
|
|
72
|
+
if name not in self._parts:
|
|
73
|
+
if not create:
|
|
74
|
+
raise KeyError("包中不存在部件 %r(如需新增请用 create=True)" % name)
|
|
75
|
+
self._order.append(name)
|
|
76
|
+
self._parts[name] = bytes(data)
|
|
77
|
+
|
|
78
|
+
def replace_text_in_part(self, name, old, new):
|
|
79
|
+
"""把一个 XML 部件当作文本做子串替换,返回替换次数。"""
|
|
80
|
+
text = self.read_part(name).decode("utf-8")
|
|
81
|
+
count = text.count(old)
|
|
82
|
+
if count:
|
|
83
|
+
self.replace_part(name, text.replace(old, new))
|
|
84
|
+
return count
|
|
85
|
+
|
|
86
|
+
# ------------------------------------------------------------------
|
|
87
|
+
# 嵌入文件(word/embeddings/)
|
|
88
|
+
# ------------------------------------------------------------------
|
|
89
|
+
|
|
90
|
+
def list_embeddings(self):
|
|
91
|
+
"""列出包内全部嵌入文件(word/embeddings/ 下的部件)。"""
|
|
92
|
+
return [p for p in self._order if p.startswith(EMBEDDINGS_DIR)]
|
|
93
|
+
|
|
94
|
+
def replace_embedded(self, name, source):
|
|
95
|
+
"""用本地文件或字节流替换包内一个嵌入文件。
|
|
96
|
+
|
|
97
|
+
name 为 zip 内路径(如 'word/embeddings/Microsoft_Excel____.xlsx'),
|
|
98
|
+
source 为文件路径或 bytes。
|
|
99
|
+
"""
|
|
100
|
+
if name not in self._parts:
|
|
101
|
+
raise KeyError("包中不存在嵌入文件 %r,现有:%r" % (name, self.list_embeddings()))
|
|
102
|
+
if isinstance(source, (bytes, bytearray)):
|
|
103
|
+
data = bytes(source)
|
|
104
|
+
else:
|
|
105
|
+
with open(source, "rb") as f:
|
|
106
|
+
data = f.read()
|
|
107
|
+
self._parts[name] = data
|
|
108
|
+
|
|
109
|
+
@contextmanager
|
|
110
|
+
def open_embedded_xlsx(self, name):
|
|
111
|
+
"""上下文管理器:直接编辑包内嵌入的 xlsx,退出时自动写回。
|
|
112
|
+
|
|
113
|
+
yield 一个 openpyxl Workbook;with 块正常结束才写回,
|
|
114
|
+
抛异常则放弃修改(包保持原样)。
|
|
115
|
+
|
|
116
|
+
例::
|
|
117
|
+
|
|
118
|
+
with pkg.open_embedded_xlsx("word/embeddings/data.xlsx") as wb:
|
|
119
|
+
ws = wb.active
|
|
120
|
+
ws["B2"] = 3.14
|
|
121
|
+
"""
|
|
122
|
+
if name not in self._parts:
|
|
123
|
+
raise KeyError("包中不存在嵌入文件 %r,现有:%r" % (name, self.list_embeddings()))
|
|
124
|
+
from openpyxl import load_workbook # 惰性导入:不用本功能就不需要 openpyxl
|
|
125
|
+
|
|
126
|
+
wb = load_workbook(io.BytesIO(self._parts[name]))
|
|
127
|
+
try:
|
|
128
|
+
yield wb
|
|
129
|
+
except Exception:
|
|
130
|
+
raise # 异常时不写回,避免半成品数据进包
|
|
131
|
+
else:
|
|
132
|
+
buf = io.BytesIO()
|
|
133
|
+
wb.save(buf)
|
|
134
|
+
self._parts[name] = buf.getvalue()
|
|
135
|
+
|
|
136
|
+
# ------------------------------------------------------------------
|
|
137
|
+
# 图表数据引用范围
|
|
138
|
+
# ------------------------------------------------------------------
|
|
139
|
+
|
|
140
|
+
def set_chart_data_rows(self, chart_part, last_row, sheet="Sheet1", start_row=2):
|
|
141
|
+
"""更新图表 XML 中数据系列的行引用范围($X$start:$X$last_row)。
|
|
142
|
+
|
|
143
|
+
对应生产场景"图表引用的行数随数据条数变化":把 chart XML 里
|
|
144
|
+
所有形如 ``Sheet1!$B$2:$B$20`` / ``Sheet1!$B$2`` 的引用改写为
|
|
145
|
+
以 last_row 结尾(或 last_row <= start_row 时收缩为单行引用)。
|
|
146
|
+
每列字母保持不变,支持多字母列($AB$)。
|
|
147
|
+
|
|
148
|
+
返回改写的引用个数。
|
|
149
|
+
"""
|
|
150
|
+
pattern = re.compile(
|
|
151
|
+
r"{sheet}!\$([A-Z]{{1,3}})\${row}(?::\$([A-Z]{{1,3}})\$(\d+))?".format(
|
|
152
|
+
sheet=re.escape(sheet), row=start_row)
|
|
153
|
+
)
|
|
154
|
+
|
|
155
|
+
def _repl(m):
|
|
156
|
+
col = m.group(1)
|
|
157
|
+
if last_row <= start_row:
|
|
158
|
+
return "{s}!${c}${r}".format(s=sheet, c=col, r=start_row)
|
|
159
|
+
return "{s}!${c}${r}:${c}${e}".format(s=sheet, c=col, r=start_row, e=last_row)
|
|
160
|
+
|
|
161
|
+
text = self.read_part(chart_part).decode("utf-8")
|
|
162
|
+
new_text, count = pattern.subn(_repl, text)
|
|
163
|
+
if count:
|
|
164
|
+
self.replace_part(chart_part, new_text)
|
|
165
|
+
return count
|
|
166
|
+
|
|
167
|
+
# ------------------------------------------------------------------
|
|
168
|
+
# 复选框打印符号修复
|
|
169
|
+
# ------------------------------------------------------------------
|
|
170
|
+
|
|
171
|
+
def fix_checkbox_glyph(self, part="word/document.xml", search="☑",
|
|
172
|
+
symbol_font="Wingdings 2", symbol_char="0052",
|
|
173
|
+
run_font="黑体", cs_font="等线"):
|
|
174
|
+
"""把 Unicode 勾选字符替换为 Wingdings 2 符号 run,保证打印正常。
|
|
175
|
+
|
|
176
|
+
背景:文档里的 "☑" 字符在部分打印流程中会显示为方框;Word 正确的
|
|
177
|
+
表达方式是一个 ``<w:sym w:font="Wingdings 2" w:char="0052"/>`` 符号。
|
|
178
|
+
本方法把 search 字符从文本 run 中切出来,替换为独立的符号 run。
|
|
179
|
+
返回替换次数。
|
|
180
|
+
"""
|
|
181
|
+
replacement = (
|
|
182
|
+
'</w:t></w:r>'
|
|
183
|
+
'<w:r><w:rPr><w:rFonts w:hint="eastAsia" w:ascii="{rf}" '
|
|
184
|
+
'w:hAnsi="{rf}" w:eastAsia="{rf}" w:cs="{cs}"/></w:rPr>'
|
|
185
|
+
'<w:sym w:font="{sf}" w:char="{sc}"/></w:r>'
|
|
186
|
+
'<w:r><w:rPr><w:rFonts w:hint="eastAsia" w:ascii="{rf}" '
|
|
187
|
+
'w:hAnsi="{rf}" w:eastAsia="{rf}" w:cs="{cs}"/></w:rPr>'
|
|
188
|
+
'<w:t xml:space="preserve">'
|
|
189
|
+
).format(rf=run_font, cs=cs_font, sf=symbol_font, sc=symbol_char)
|
|
190
|
+
return self.replace_text_in_part(part, search, replacement)
|
|
191
|
+
|
|
192
|
+
# ------------------------------------------------------------------
|
|
193
|
+
# 保存
|
|
194
|
+
# ------------------------------------------------------------------
|
|
195
|
+
|
|
196
|
+
def save(self, path):
|
|
197
|
+
"""写回磁盘。未修改的部件按原始顺序、原始字节写入(ZIP_DEFLATED)。"""
|
|
198
|
+
with zipfile.ZipFile(path, "w", zipfile.ZIP_DEFLATED) as zf:
|
|
199
|
+
for name in self._order:
|
|
200
|
+
zf.writestr(name, self._parts[name])
|
|
201
|
+
|
|
202
|
+
# ------------------------------------------------------------------
|
|
203
|
+
# 内部
|
|
204
|
+
# ------------------------------------------------------------------
|
|
205
|
+
|
|
206
|
+
def _require_part(self, name):
|
|
207
|
+
if name not in self._parts:
|
|
208
|
+
raise KeyError("包中不存在部件 %r" % name)
|
|
209
|
+
|
|
210
|
+
|
|
211
|
+
def set_table_range(ws, ref):
|
|
212
|
+
"""修改工作表中 Excel 表格(ListObject)的数据范围。
|
|
213
|
+
|
|
214
|
+
在 open_embedded_xlsx 的 with 块内配合使用::
|
|
215
|
+
|
|
216
|
+
with pkg.open_embedded_xlsx(...) as wb:
|
|
217
|
+
ws = wb.active
|
|
218
|
+
ws["A2"] = "1月"
|
|
219
|
+
set_table_range(ws, "A1:C8")
|
|
220
|
+
|
|
221
|
+
会更新该工作表上的全部表格,返回表格数量。
|
|
222
|
+
"""
|
|
223
|
+
tables = list(getattr(ws, "tables", {}).values())
|
|
224
|
+
if not tables:
|
|
225
|
+
raise ValueError("工作表 %r 中没有 Excel 表格(ListObject),无法设置范围" % ws.title)
|
|
226
|
+
for table in tables:
|
|
227
|
+
table.ref = ref
|
|
228
|
+
return len(tables)
|
|
@@ -0,0 +1,203 @@
|
|
|
1
|
+
Metadata-Version: 2.4
|
|
2
|
+
Name: docxsurgeon
|
|
3
|
+
Version: 0.1.0
|
|
4
|
+
Summary: 对 docx 做 zip 包级手术:替换/编辑嵌入的 xlsx、更新图表数据范围、修复复选框打印符号
|
|
5
|
+
Author: txxcat
|
|
6
|
+
License: MIT
|
|
7
|
+
Project-URL: Homepage, https://github.com/txxcat/docxsurgeon
|
|
8
|
+
Keywords: docx,ooxml,embedded,excel,chart,word,surgery
|
|
9
|
+
Classifier: Development Status :: 3 - Alpha
|
|
10
|
+
Classifier: Intended Audience :: Developers
|
|
11
|
+
Classifier: License :: OSI Approved :: MIT License
|
|
12
|
+
Classifier: Operating System :: OS Independent
|
|
13
|
+
Classifier: Programming Language :: Python :: 3
|
|
14
|
+
Classifier: Programming Language :: Python :: 3.9
|
|
15
|
+
Classifier: Programming Language :: Python :: 3.10
|
|
16
|
+
Classifier: Programming Language :: Python :: 3.11
|
|
17
|
+
Classifier: Programming Language :: Python :: 3.12
|
|
18
|
+
Classifier: Programming Language :: Python :: 3.13
|
|
19
|
+
Classifier: Topic :: Office/Business :: Office Suites
|
|
20
|
+
Classifier: Topic :: Software Development :: Libraries :: Python Modules
|
|
21
|
+
Requires-Python: >=3.9
|
|
22
|
+
Description-Content-Type: text/markdown
|
|
23
|
+
License-File: LICENSE
|
|
24
|
+
Requires-Dist: openpyxl>=3.0
|
|
25
|
+
Dynamic: license-file
|
|
26
|
+
|
|
27
|
+
# docxsurgeon
|
|
28
|
+
|
|
29
|
+
[](https://github.com/txxcat/docxsurgeon/actions/workflows/ci.yml)
|
|
30
|
+
[](LICENSE)
|
|
31
|
+
[](https://www.python.org)
|
|
32
|
+
[](#roadmap)
|
|
33
|
+
|
|
34
|
+
**Operate on the *inside* of .docx files — replace embedded Excel, update chart ranges, fix checkbox glyphs. All in memory.**
|
|
35
|
+
|
|
36
|
+
**对 docx 做 zip 包级"手术":替换/编辑嵌入的 xlsx、更新图表数据引用范围、修复复选框打印符号。全程内存操作,不落地临时目录。**
|
|
37
|
+
|
|
38
|
+
## 为什么需要它
|
|
39
|
+
|
|
40
|
+
OOXML 文件(docx/xlsx/pptx)本质是 zip 包,里面藏着 `word/embeddings/*.xlsx`(嵌入的 Excel)、`word/charts/chart*.xml`(图表定义)、`xl/tables/*.xml`(Excel 表格范围)。这些**包内结构**在主流库里无人认领:
|
|
41
|
+
|
|
42
|
+
| 需求 | python-docx | docxtpl | openpyxl | docxsurgeon |
|
|
43
|
+
|---|---|---|---|---|
|
|
44
|
+
| 替换 docx 内嵌入的 xlsx | ❌ | ⚠️ 需走 Jinja2 render 流程 | ❌(只认独立文件) | ✅ |
|
|
45
|
+
| **编辑**嵌入 xlsx 的单元格 | ❌ | ❌ | ❌ | ✅ 内存直改,退出自动写回 |
|
|
46
|
+
| 改嵌入 xlsx 里 Excel 表格范围 | ❌ | ❌ | ❌ | ✅ |
|
|
47
|
+
| 更新图表 XML 数据引用行数 | ❌ | ❌ | ❌ | ✅ |
|
|
48
|
+
| Unicode ☑ → Wingdings 符号 run(打印修复) | ❌ | ❌ | ❌ | ✅ |
|
|
49
|
+
| 文档模板变量填充 | ❌ | ✅ | — | ❌(不抢 docxtpl 的活) |
|
|
50
|
+
| 文档模型读写(段落/表格) | ✅ | ✅ | — | ❌(不抢 python-docx 的活) |
|
|
51
|
+
|
|
52
|
+
一句话定位:**python-docx 管文档模型,docxtpl 管模板填充,docxsurgeon 管"已有文档的包级修改"。**
|
|
53
|
+
|
|
54
|
+
核心逻辑提取自一套在广电播控机房生产环境运行多年的报表工具,经过真实 Word / WPS 文档的长期验证。
|
|
55
|
+
|
|
56
|
+
## 安装
|
|
57
|
+
|
|
58
|
+
要求 **Python >= 3.9**。
|
|
59
|
+
|
|
60
|
+
**从 PyPI(待发布)**
|
|
61
|
+
|
|
62
|
+
> 本项目尚未在 PyPI 正式发布,0.1.0 上线后会启用下面的命令。当前请使用下方的源码安装方式。
|
|
63
|
+
|
|
64
|
+
```bash
|
|
65
|
+
pip install docxsurgeon
|
|
66
|
+
```
|
|
67
|
+
|
|
68
|
+
**从源码安装(当前阶段推荐)**
|
|
69
|
+
|
|
70
|
+
```bash
|
|
71
|
+
pip install "git+https://github.com/txxcat/docxsurgeon.git"
|
|
72
|
+
```
|
|
73
|
+
|
|
74
|
+
**开发模式**
|
|
75
|
+
|
|
76
|
+
```bash
|
|
77
|
+
git clone https://github.com/txxcat/docxsurgeon
|
|
78
|
+
cd docxsurgeon
|
|
79
|
+
pip install -e .
|
|
80
|
+
```
|
|
81
|
+
|
|
82
|
+
依赖说明:运行时仅依赖 `openpyxl>=3.0`(编辑嵌入 xlsx 时用到,已随安装自动拉取)。其余功能(部件读写、图表范围改写、符号修复)不依赖任何第三方库。
|
|
83
|
+
|
|
84
|
+
## 快速上手
|
|
85
|
+
|
|
86
|
+
```python
|
|
87
|
+
from docxsurgeon import DocxPackage, set_table_range
|
|
88
|
+
|
|
89
|
+
pkg = DocxPackage("周报模板.docx")
|
|
90
|
+
|
|
91
|
+
# 1) 直接编辑包内嵌入的 Excel(内存中,无临时目录)
|
|
92
|
+
with pkg.open_embedded_xlsx("word/embeddings/Microsoft_Excel____1.xlsx") as wb:
|
|
93
|
+
ws = wb.active
|
|
94
|
+
ws["A2"] = "9月"
|
|
95
|
+
ws["B2"] = 12
|
|
96
|
+
set_table_range(ws, "A1:C8") # 同步 Excel 表格范围
|
|
97
|
+
|
|
98
|
+
# 2) 图表引用行数随数据条数伸缩(Sheet1!$B$2:$B$20 -> $B$2:$B$8)
|
|
99
|
+
pkg.set_chart_data_rows("word/charts/chart2.xml", last_row=8)
|
|
100
|
+
|
|
101
|
+
# 3) 替换嵌入文件本体
|
|
102
|
+
pkg.replace_embedded("word/embeddings/Microsoft_Word_Document3.docx", "新附件.docx")
|
|
103
|
+
|
|
104
|
+
# 4) 修复 ☑ 打印乱码:替换为 Wingdings 2 符号 run
|
|
105
|
+
pkg.fix_checkbox_glyph()
|
|
106
|
+
|
|
107
|
+
pkg.save("输出.docx") # 未修改的部件字节级原样保留
|
|
108
|
+
```
|
|
109
|
+
|
|
110
|
+
## 进阶示例
|
|
111
|
+
|
|
112
|
+
**探索文档结构** —— 动刀前先看看包里有什么:
|
|
113
|
+
|
|
114
|
+
```python
|
|
115
|
+
pkg = DocxPackage("周报模板.docx")
|
|
116
|
+
print(pkg.list_parts()) # 全部部件,按原始 zip 顺序
|
|
117
|
+
print(pkg.list_embeddings()) # 只列 word/embeddings/ 下的嵌入文件
|
|
118
|
+
print(pkg.has_part("word/document.xml")) # True
|
|
119
|
+
```
|
|
120
|
+
|
|
121
|
+
**部件内文本替换** —— 把某个 XML 部件当纯文本做子串替换,返回替换次数:
|
|
122
|
+
|
|
123
|
+
```python
|
|
124
|
+
n = pkg.replace_text_in_part(
|
|
125
|
+
"word/charts/chart2.xml",
|
|
126
|
+
"Sheet1!$B$2:$B$20",
|
|
127
|
+
"Sheet1!$B$2:$B$8",
|
|
128
|
+
)
|
|
129
|
+
print(f"改了 {n} 处")
|
|
130
|
+
```
|
|
131
|
+
|
|
132
|
+
**失败安全** —— `open_embedded_xlsx` 的 with 块里若抛异常,本次对嵌入簿的修改会被整体丢弃,包保持原样:
|
|
133
|
+
|
|
134
|
+
```python
|
|
135
|
+
try:
|
|
136
|
+
with pkg.open_embedded_xlsx("word/embeddings/data.xlsx") as wb:
|
|
137
|
+
wb.active["A1"] = "新值"
|
|
138
|
+
raise RuntimeError("模拟中途出错")
|
|
139
|
+
except RuntimeError:
|
|
140
|
+
pass
|
|
141
|
+
# 此时嵌入簿未被破坏,可安全地改做别的处理或直接 save()
|
|
142
|
+
```
|
|
143
|
+
|
|
144
|
+
## API
|
|
145
|
+
|
|
146
|
+
### `DocxPackage(source)`
|
|
147
|
+
|
|
148
|
+
`source` 为文件路径或二进制文件对象。
|
|
149
|
+
|
|
150
|
+
| 方法 | 说明 |
|
|
151
|
+
|---|---|
|
|
152
|
+
| `list_parts()` / `has_part(name)` / `read_part(name)` | zip 部件级读取 |
|
|
153
|
+
| `replace_part(name, data, create=False)` | 整体替换部件(bytes/str) |
|
|
154
|
+
| `replace_text_in_part(name, old, new)` | 部件内文本替换,返回次数 |
|
|
155
|
+
| `list_embeddings()` | 列出 `word/embeddings/` 下的全部嵌入文件 |
|
|
156
|
+
| `replace_embedded(name, source)` | 用本地文件或字节替换嵌入文件 |
|
|
157
|
+
| `open_embedded_xlsx(name)` | **上下文管理器**:yield openpyxl Workbook,正常退出自动写回,异常退出放弃修改 |
|
|
158
|
+
| `set_chart_data_rows(chart_part, last_row, sheet="Sheet1", start_row=2)` | 改写图表数据引用范围,返回改写数 |
|
|
159
|
+
| `fix_checkbox_glyph(part, search="☑", symbol_font="Wingdings 2", symbol_char="0052", ...)` | Unicode 勾选符 → 符号 run |
|
|
160
|
+
| `save(path)` | 写回磁盘 |
|
|
161
|
+
|
|
162
|
+
### 模块函数
|
|
163
|
+
|
|
164
|
+
- `set_table_range(ws, ref)` —— 在 `open_embedded_xlsx` 块内修改 Excel 表格(ListObject)范围。
|
|
165
|
+
- `__version__` —— 当前版本号(与 `package.py` 单一来源一致)。
|
|
166
|
+
|
|
167
|
+
## 设计要点
|
|
168
|
+
|
|
169
|
+
- **全内存**:部件读写走 `BytesIO`,没有临时目录、没有 `os.chdir`、没有竞态清理。
|
|
170
|
+
- **保真**:`save()` 按原 zip 顺序写入;未修改的部件字节级不变(测试保证)。
|
|
171
|
+
- **失败安全**:`open_embedded_xlsx` 中抛异常则放弃本次修改,包保持原样。
|
|
172
|
+
- **惰性依赖**:只有用到 `open_embedded_xlsx` 才需要 openpyxl。
|
|
173
|
+
|
|
174
|
+
## 已知边界
|
|
175
|
+
|
|
176
|
+
- 图表的缓存数据(`c:strCache` / `c:numCache`)不会同步改写——Word 打开时会从嵌入工作簿刷新。若你的场景要求"不刷新也立即显示新值",参考 roadmap。
|
|
177
|
+
- 嵌入的 PDF/OLE 对象(`oleObjectNNN.bin`)不支持替换(OLE 是编码格式)。
|
|
178
|
+
- `open_embedded_xlsx` 依赖 openpyxl 往返,openpyxl 不建模的 xlsx 特性会丢失(与直接用 openpyxl 编辑同一文件的风险一致)。
|
|
179
|
+
|
|
180
|
+
## Roadmap
|
|
181
|
+
|
|
182
|
+
- [ ] 上传 PyPI 正式版(当前未发布,`pip install` 需走源码)
|
|
183
|
+
- [ ] 图表缓存数据(strCache/numCache)双写
|
|
184
|
+
- [ ] 嵌入 xlsx 的图表部件(`word/charts` 引用嵌入簿的完整链路更新)
|
|
185
|
+
- [ ] pptx / xlsx 宿主包的一等支持(当前 API 以 docx 命名,机制通用)
|
|
186
|
+
- [ ] CLI:`python -m docxsurgeon list/extract/replace ...`
|
|
187
|
+
|
|
188
|
+
## 开发 & 贡献
|
|
189
|
+
|
|
190
|
+
```bash
|
|
191
|
+
git clone https://github.com/txxcat/docxsurgeon
|
|
192
|
+
cd docxsurgeon
|
|
193
|
+
pip install -e . pytest
|
|
194
|
+
pytest
|
|
195
|
+
```
|
|
196
|
+
|
|
197
|
+
仓库通过 GitHub Actions 跑 CI(见 `.github/workflows/ci.yml`,每次 push/PR 自动跑测试)。
|
|
198
|
+
|
|
199
|
+
想参与贡献?请先阅读 [CONTRIBUTING.md](CONTRIBUTING.md)。
|
|
200
|
+
|
|
201
|
+
## License
|
|
202
|
+
|
|
203
|
+
[MIT](LICENSE)
|
|
@@ -0,0 +1,13 @@
|
|
|
1
|
+
LICENSE
|
|
2
|
+
README.md
|
|
3
|
+
pyproject.toml
|
|
4
|
+
src/docxsurgeon/__init__.py
|
|
5
|
+
src/docxsurgeon/package.py
|
|
6
|
+
src/docxsurgeon.egg-info/PKG-INFO
|
|
7
|
+
src/docxsurgeon.egg-info/SOURCES.txt
|
|
8
|
+
src/docxsurgeon.egg-info/dependency_links.txt
|
|
9
|
+
src/docxsurgeon.egg-info/requires.txt
|
|
10
|
+
src/docxsurgeon.egg-info/top_level.txt
|
|
11
|
+
tests/test_chart_and_symbol.py
|
|
12
|
+
tests/test_embedded.py
|
|
13
|
+
tests/test_package.py
|
|
@@ -0,0 +1 @@
|
|
|
1
|
+
|
|
@@ -0,0 +1 @@
|
|
|
1
|
+
openpyxl>=3.0
|
|
@@ -0,0 +1 @@
|
|
|
1
|
+
docxsurgeon
|
|
@@ -0,0 +1,89 @@
|
|
|
1
|
+
"""图表引用范围与复选框符号测试。"""
|
|
2
|
+
import zipfile
|
|
3
|
+
|
|
4
|
+
import pytest
|
|
5
|
+
|
|
6
|
+
from docxsurgeon import DocxPackage
|
|
7
|
+
|
|
8
|
+
CHART = "word/charts/chart1.xml"
|
|
9
|
+
|
|
10
|
+
CHART_XML_MULTI_COLUMN = (
|
|
11
|
+
'<?xml version="1.0" encoding="UTF-8" standalone="yes"?>\n'
|
|
12
|
+
'<c:chartSpace xmlns:c="http://schemas.openxmlformats.org/drawingml/2006/chart">'
|
|
13
|
+
'<c:chart><c:plotArea><c:lineChart><c:ser>'
|
|
14
|
+
'<c:val><c:numRef><c:f>Sheet1!$AB$2:$AB$20</c:f></c:numRef></c:val>'
|
|
15
|
+
'<c:val><c:numRef><c:f>Sheet1!$AC$2:$AC$20</c:f></c:numRef></c:val>'
|
|
16
|
+
'</c:ser></c:lineChart></c:plotArea></c:chart></c:chartSpace>'
|
|
17
|
+
)
|
|
18
|
+
|
|
19
|
+
|
|
20
|
+
def test_set_chart_data_rows_basic(pkg):
|
|
21
|
+
n = pkg.set_chart_data_rows(CHART, last_row=8)
|
|
22
|
+
assert n == 2 # cat + val 两处引用
|
|
23
|
+
text = pkg.read_part(CHART).decode("utf-8")
|
|
24
|
+
assert "Sheet1!$A$2:$A$8" in text
|
|
25
|
+
assert "Sheet1!$B$2:$B$8" in text
|
|
26
|
+
assert "Sheet1!$A$2:$A$20" not in text
|
|
27
|
+
assert "Sheet1!$B$2:$B$20" not in text
|
|
28
|
+
# 系列名引用($B$1,单行)不受影响
|
|
29
|
+
assert "Sheet1!$B$1" in text
|
|
30
|
+
|
|
31
|
+
|
|
32
|
+
def test_set_chart_data_rows_shrink_to_single(pkg):
|
|
33
|
+
n = pkg.set_chart_data_rows(CHART, last_row=2)
|
|
34
|
+
assert n == 2
|
|
35
|
+
text = pkg.read_part(CHART).decode("utf-8")
|
|
36
|
+
assert "Sheet1!$A$2" in text
|
|
37
|
+
assert "$A$2:$A$" not in text, "last_row<=start_row 时应收缩为单行引用"
|
|
38
|
+
|
|
39
|
+
|
|
40
|
+
def test_set_chart_data_rows_multi_letter_columns(tmp_path):
|
|
41
|
+
path = tmp_path / "multi.docx"
|
|
42
|
+
with zipfile.ZipFile(path, "w") as zf:
|
|
43
|
+
zf.writestr("word/charts/chart1.xml", CHART_XML_MULTI_COLUMN)
|
|
44
|
+
pkg = DocxPackage(path)
|
|
45
|
+
n = pkg.set_chart_data_rows("word/charts/chart1.xml", last_row=12)
|
|
46
|
+
assert n == 2
|
|
47
|
+
text = pkg.read_part("word/charts/chart1.xml").decode("utf-8")
|
|
48
|
+
assert "Sheet1!$AB$2:$AB$12" in text
|
|
49
|
+
assert "Sheet1!$AC$2:$AC$12" in text
|
|
50
|
+
|
|
51
|
+
|
|
52
|
+
def test_set_chart_data_rows_other_sheet_untouched(tmp_path):
|
|
53
|
+
path = tmp_path / "sheet.docx"
|
|
54
|
+
xml = ('<c:chartSpace xmlns:c="http://x">'
|
|
55
|
+
'<c:f>Sheet1!$B$2:$B$20</c:f>'
|
|
56
|
+
'<c:f>数据!$B$2:$B$20</c:f>'
|
|
57
|
+
'</c:chartSpace>')
|
|
58
|
+
with zipfile.ZipFile(path, "w") as zf:
|
|
59
|
+
zf.writestr("word/charts/chart1.xml", xml)
|
|
60
|
+
pkg = DocxPackage(path)
|
|
61
|
+
n = pkg.set_chart_data_rows("word/charts/chart1.xml", last_row=6)
|
|
62
|
+
assert n == 1, "只应改写指定 sheet 的引用"
|
|
63
|
+
text = pkg.read_part("word/charts/chart1.xml").decode("utf-8")
|
|
64
|
+
assert "Sheet1!$B$2:$B$6" in text
|
|
65
|
+
assert "数据!$B$2:$B$20" in text
|
|
66
|
+
|
|
67
|
+
|
|
68
|
+
def test_set_chart_data_rows_no_match(pkg):
|
|
69
|
+
n = pkg.set_chart_data_rows("word/charts/chart1.xml", last_row=8, sheet="不存在的表")
|
|
70
|
+
assert n == 0
|
|
71
|
+
|
|
72
|
+
|
|
73
|
+
def test_fix_checkbox_glyph(pkg):
|
|
74
|
+
n = pkg.fix_checkbox_glyph()
|
|
75
|
+
assert n == 2, "文档中有两个 ☑"
|
|
76
|
+
text = pkg.read_part("word/document.xml").decode("utf-8")
|
|
77
|
+
assert "☑" not in text
|
|
78
|
+
assert '<w:sym w:font="Wingdings 2" w:char="0052"/>' in text
|
|
79
|
+
# 符号 run 前后正确闭合了文本 run
|
|
80
|
+
assert text.count("</w:t></w:r><w:r><w:rPr>") == 2
|
|
81
|
+
assert text.count('<w:t xml:space="preserve">') == 2
|
|
82
|
+
# 没有动其他部件
|
|
83
|
+
assert pkg.has_part("word/embeddings/data.xlsx")
|
|
84
|
+
|
|
85
|
+
|
|
86
|
+
def test_fix_checkbox_glyph_custom_symbol(pkg):
|
|
87
|
+
pkg.fix_checkbox_glyph(symbol_font="Wingdings", symbol_char="00FC")
|
|
88
|
+
text = pkg.read_part("word/document.xml").decode("utf-8")
|
|
89
|
+
assert 'w:font="Wingdings" w:char="00FC"' in text
|
|
@@ -0,0 +1,75 @@
|
|
|
1
|
+
"""嵌入 xlsx 编辑测试。"""
|
|
2
|
+
import io
|
|
3
|
+
import zipfile
|
|
4
|
+
|
|
5
|
+
import pytest
|
|
6
|
+
from openpyxl import load_workbook
|
|
7
|
+
|
|
8
|
+
from docxsurgeon import set_table_range
|
|
9
|
+
|
|
10
|
+
EMBED = "word/embeddings/data.xlsx"
|
|
11
|
+
|
|
12
|
+
|
|
13
|
+
def _read_table_ref(pkg):
|
|
14
|
+
"""从包内的嵌入 xlsx 字节里直接读 xl/tables/table1.xml 的 ref 属性。"""
|
|
15
|
+
import re
|
|
16
|
+
with zipfile.ZipFile(io.BytesIO(pkg.read_part(EMBED))) as zf:
|
|
17
|
+
xml = zf.read("xl/tables/table1.xml").decode("utf-8")
|
|
18
|
+
m = re.search(r'ref="([^"]+)"', xml)
|
|
19
|
+
return m.group(1)
|
|
20
|
+
|
|
21
|
+
|
|
22
|
+
def test_open_embedded_xlsx_write(pkg):
|
|
23
|
+
with pkg.open_embedded_xlsx(EMBED) as wb:
|
|
24
|
+
ws = wb.active
|
|
25
|
+
ws["A2"] = "9月"
|
|
26
|
+
ws["B2"] = 42
|
|
27
|
+
set_table_range(ws, "A1:B8")
|
|
28
|
+
|
|
29
|
+
# 单元格修改已写回
|
|
30
|
+
with pkg.open_embedded_xlsx(EMBED) as wb:
|
|
31
|
+
ws = wb.active
|
|
32
|
+
assert ws["A2"].value == "9月"
|
|
33
|
+
assert ws["B2"].value == 42
|
|
34
|
+
# 表格范围已写回(直接检查嵌套 zip 的 XML)
|
|
35
|
+
assert _read_table_ref(pkg) == "A1:B8"
|
|
36
|
+
|
|
37
|
+
|
|
38
|
+
def test_open_embedded_xlsx_exception_no_write(pkg):
|
|
39
|
+
original = pkg.read_part(EMBED)
|
|
40
|
+
with pytest.raises(RuntimeError):
|
|
41
|
+
with pkg.open_embedded_xlsx(EMBED) as wb:
|
|
42
|
+
wb.active["A2"] = "不该写进去"
|
|
43
|
+
raise RuntimeError("boom")
|
|
44
|
+
assert pkg.read_part(EMBED) == original, "异常退出时不应写回修改"
|
|
45
|
+
|
|
46
|
+
|
|
47
|
+
def test_set_table_range_multiple_tables(pkg):
|
|
48
|
+
with pkg.open_embedded_xlsx(EMBED) as wb:
|
|
49
|
+
ws = wb.active
|
|
50
|
+
n = set_table_range(ws, "A1:B10")
|
|
51
|
+
assert n == 1
|
|
52
|
+
|
|
53
|
+
|
|
54
|
+
def test_set_table_range_no_table_raises():
|
|
55
|
+
from openpyxl import Workbook
|
|
56
|
+
ws = Workbook().active
|
|
57
|
+
with pytest.raises(ValueError):
|
|
58
|
+
set_table_range(ws, "A1:B5")
|
|
59
|
+
|
|
60
|
+
|
|
61
|
+
def test_open_embedded_xlsx_unknown_raises(pkg):
|
|
62
|
+
with pytest.raises(KeyError):
|
|
63
|
+
with pkg.open_embedded_xlsx("word/embeddings/nothere.xlsx"):
|
|
64
|
+
pass
|
|
65
|
+
|
|
66
|
+
|
|
67
|
+
def test_roundtrip_through_openpyxl_keeps_table(pkg, tmp_path):
|
|
68
|
+
"""openpyxl 往返后嵌入文件仍是合法 xlsx 且表格仍在。"""
|
|
69
|
+
with pkg.open_embedded_xlsx(EMBED) as wb:
|
|
70
|
+
wb.active["C1"] = "新增列"
|
|
71
|
+
out = tmp_path / "roundtrip.xlsx"
|
|
72
|
+
out.write_bytes(pkg.read_part(EMBED))
|
|
73
|
+
wb2 = load_workbook(out)
|
|
74
|
+
assert wb2.active["C1"].value == "新增列"
|
|
75
|
+
assert len(wb2.active.tables) == 1, "表格在 openpyxl 往返后丢失"
|
|
@@ -0,0 +1,83 @@
|
|
|
1
|
+
"""部件级操作与保存保真测试。"""
|
|
2
|
+
import zipfile
|
|
3
|
+
|
|
4
|
+
from docxsurgeon import DocxPackage
|
|
5
|
+
|
|
6
|
+
|
|
7
|
+
def test_unmodified_parts_byte_identical(docx_path, tmp_path):
|
|
8
|
+
"""保存后,未修改的部件必须字节级原样保留。"""
|
|
9
|
+
pkg = DocxPackage(docx_path)
|
|
10
|
+
out = tmp_path / "out.docx"
|
|
11
|
+
pkg.save(out)
|
|
12
|
+
|
|
13
|
+
with zipfile.ZipFile(docx_path) as a, zipfile.ZipFile(out) as b:
|
|
14
|
+
names_a = a.namelist()
|
|
15
|
+
assert b.namelist() == names_a, "部件顺序或集合发生变化"
|
|
16
|
+
for name in names_a:
|
|
17
|
+
assert a.read(name) == b.read(name), "部件 %r 内容被意外改动" % name
|
|
18
|
+
|
|
19
|
+
|
|
20
|
+
def test_list_and_read_parts(pkg):
|
|
21
|
+
parts = pkg.list_parts()
|
|
22
|
+
assert "word/document.xml" in parts
|
|
23
|
+
assert "word/embeddings/data.xlsx" in parts
|
|
24
|
+
assert pkg.has_part("word/charts/chart1.xml")
|
|
25
|
+
assert not pkg.has_part("word/nothing.xml")
|
|
26
|
+
assert pkg.read_part("[Content_Types].xml") == b"<Types/>"
|
|
27
|
+
|
|
28
|
+
|
|
29
|
+
def test_replace_part(pkg):
|
|
30
|
+
pkg.replace_part("word/document.xml", "<new/>")
|
|
31
|
+
assert pkg.read_part("word/document.xml") == b"<new/>"
|
|
32
|
+
# 其他部件不受影响
|
|
33
|
+
assert pkg.has_part("word/charts/chart1.xml")
|
|
34
|
+
|
|
35
|
+
|
|
36
|
+
def test_replace_part_unknown_raises(pkg):
|
|
37
|
+
import pytest
|
|
38
|
+
with pytest.raises(KeyError):
|
|
39
|
+
pkg.replace_part("word/ghost.xml", b"x")
|
|
40
|
+
# create=True 允许新增,且追加在包尾
|
|
41
|
+
pkg.replace_part("word/ghost.xml", b"x", create=True)
|
|
42
|
+
assert pkg.list_parts()[-1] == "word/ghost.xml"
|
|
43
|
+
|
|
44
|
+
|
|
45
|
+
def test_read_part_unknown_raises(pkg):
|
|
46
|
+
import pytest
|
|
47
|
+
with pytest.raises(KeyError):
|
|
48
|
+
pkg.read_part("word/ghost.xml")
|
|
49
|
+
|
|
50
|
+
|
|
51
|
+
def test_replace_text_in_part_count(pkg):
|
|
52
|
+
n = pkg.replace_text_in_part("word/document.xml", "☑", "OK")
|
|
53
|
+
assert n == 2
|
|
54
|
+
assert b"\xe2\x98\x91" not in pkg.read_part("word/document.xml")
|
|
55
|
+
assert pkg.read_part("word/document.xml").decode("utf-8").count("OK") == 2
|
|
56
|
+
|
|
57
|
+
|
|
58
|
+
def test_list_embeddings(pkg):
|
|
59
|
+
assert pkg.list_embeddings() == ["word/embeddings/data.xlsx"]
|
|
60
|
+
|
|
61
|
+
|
|
62
|
+
def test_replace_embedded_with_file(pkg, tmp_path):
|
|
63
|
+
new_xlsx = tmp_path / "new.xlsx"
|
|
64
|
+
new_xlsx.write_bytes(b"FAKE_XLSX_BYTES")
|
|
65
|
+
pkg.replace_embedded("word/embeddings/data.xlsx", str(new_xlsx))
|
|
66
|
+
assert pkg.read_part("word/embeddings/data.xlsx") == b"FAKE_XLSX_BYTES"
|
|
67
|
+
|
|
68
|
+
|
|
69
|
+
def test_replace_embedded_with_bytes(pkg):
|
|
70
|
+
pkg.replace_embedded("word/embeddings/data.xlsx", b"BYTES")
|
|
71
|
+
assert pkg.read_part("word/embeddings/data.xlsx") == b"BYTES"
|
|
72
|
+
|
|
73
|
+
|
|
74
|
+
def test_replace_embedded_unknown_raises(pkg):
|
|
75
|
+
import pytest
|
|
76
|
+
with pytest.raises(KeyError):
|
|
77
|
+
pkg.replace_embedded("word/embeddings/nothere.xlsx", b"x")
|
|
78
|
+
|
|
79
|
+
|
|
80
|
+
def test_load_from_file_object(docx_path):
|
|
81
|
+
with open(docx_path, "rb") as f:
|
|
82
|
+
pkg = DocxPackage(f)
|
|
83
|
+
assert pkg.has_part("word/document.xml")
|