epub-extended 0.1.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 Taylor Ren
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
@@ -0,0 +1,12 @@
1
+ # The wheel ships only the package (see [tool.setuptools] in pyproject.toml).
2
+ # The sdist is the source of the project, so it carries the tests, the dev
3
+ # tools, and the docs the README points at.
4
+ include README.md
5
+ include SPEC.md
6
+ include PITFALLS.md
7
+ include RELEASING.md
8
+
9
+ recursive-include tests *.py
10
+ recursive-include tools *.py
11
+
12
+ global-exclude __pycache__ *.py[cod]
@@ -0,0 +1,363 @@
1
+ # epubx — Pitfalls
2
+
3
+ Every entry here is a bug that **epubx actually shipped**, found by running
4
+ against real books. Each one produced plausible-looking output — no exception,
5
+ no warning — and silently lost or misreported content. Synthetic fixtures did
6
+ not catch any of them.
7
+
8
+ Read this before changing the parser. Most of these are ways to be *quietly*
9
+ wrong, which is the only failure mode that matters in a library whose job is to
10
+ report what a document contains.
11
+
12
+ **Evidence base:** 300 EPUB files (1.39 GB, mostly Chinese, from a Calibre
13
+ library). Current state: **0 uncaught exceptions**, 661,108 paragraphs,
14
+ 7,900/7,905 footnote references resolved, 4 books named image-only.
15
+
16
+ ---
17
+
18
+ ## The rule that explains all of them
19
+
20
+ > **A block walker that finds no block has not found no content.**
21
+
22
+ Recursing into a wrapper and emitting nothing looks identical to a wrapper that
23
+ genuinely is empty. Every data-loss bug below is that confusion. When a
24
+ container yields zero blocks, ask what it *actually* held before concluding it
25
+ was empty — the text is usually still there, sitting somewhere you did not look.
26
+
27
+ Corollary, learned the hard way twice: **never infer a book's nature from an
28
+ aggregate.** Sampled chapters, summed character counts and averaged ratios each
29
+ produced a confident, wrong verdict on a real book. Look at one document's raw
30
+ markup before believing any statistic.
31
+
32
+ ---
33
+
34
+ ## 1. Paragraphs that are not `<p>`
35
+
36
+ **The single largest source of lost text.**
37
+
38
+ `<p>` is not a requirement of XHTML-in-practice. Real books emit paragraphs as:
39
+
40
+ ```html
41
+ <!-- Calibre convention: a class, and no <p> anywhere in the file -->
42
+ <div class="p-indent"><span>Body text…</span></div>
43
+ ```
44
+
45
+ ```html
46
+ <!-- Chinese novels: prose set loose, separated by <br/> -->
47
+ <div class="calibre1">
48
+ <h3>第九部:神秘敵人</h3><br/>
49
+ <br/>
50
+   黃俊和兩個大漢,跟在我們背後…<br/>
51
+ <br/>
52
+   「死神」?不可能的…
53
+ </div>
54
+ ```
55
+
56
+ Two separate traps:
57
+
58
+ - **The Calibre class.** Treat every `<div>` as structure and *Killing Lincoln*
59
+ (79 chapters, 511,617 characters) parsed to **zero**. Honour
60
+ `p-indent`/`p-br`/`p-blanc`/`p-continuance`.
61
+ - **Prose in the tails.** The second form's text is in the **`.tail` of each
62
+ `<br/>`**, not in the div's own `.text` (which is indentation whitespace).
63
+ Checking only `.text` finds nothing. ~130 readable books parsed to ~465
64
+ characters each; they now yield their full text.
65
+
66
+ **And the trap in the fix.** The obvious generalisation — "a div with no block
67
+ children is a paragraph" — is wrong: it swallows `<body>` and `<section>` and
68
+ collapses an entire chapter into one block (it broke 10 tests). And "contains a
69
+ block child" is *also* wrong, because the wrapper above legitimately contains an
70
+ `<h3>`. The rule that holds: a block **containing the prose** disqualifies;
71
+ inline markup and headings alongside loose prose do not.
72
+
73
+ ## 2. lxml's HTML parser reads undeclared files as latin-1
74
+
75
+ A document with **no charset declaration at all** — valid UTF-8 — came out as
76
+ `Whatâ\x80\x99s next`. lxml sniffs, fails to find a declaration, falls back to
77
+ latin-1, and the damage is invisible until you read the text.
78
+
79
+ **The fix that does not work:** decoding the bytes yourself and re-encoding as
80
+ UTF-8. lxml re-sniffs and guesses latin-1 again. You must pass the encoding to
81
+ the parser:
82
+
83
+ ```python
84
+ html.HTMLParser(recover=True, encoding="utf-8") # this is the load-bearing part
85
+ ```
86
+
87
+ ## 3. `xlink:href` is not namespaced by the HTML parser
88
+
89
+ An SVG-wrapped cover reported **zero images**. The XML parser expands
90
+ `xlink:href` to `{http://www.w3.org/1999/xlink}href`; the **HTML parser leaves
91
+ the literal string `"xlink:href"`**. Reading only the Clark-notation form finds
92
+ nothing, with no error. Accept both.
93
+
94
+ Corollary: books wrap covers in `<svg><image>` rather than `<img>`. Handle SVG
95
+ images or whole pages vanish.
96
+
97
+ ## 4. `opf:file-as` is namespaced
98
+
99
+ `<dc:creator opf:file-as="Doe, Jane">` parses to
100
+ `{http://www.idpf.org/2007/opf}file-as`. `element.get("file-as")` returns `None`,
101
+ silently, for every creator in the book.
102
+
103
+ ## 5. `encryption.xml` is not DRM
104
+
105
+ Font obfuscation XOR-encodes `.ttf` files to discourage extraction and **leaves
106
+ the text in the clear**. Two corpus books do exactly this — one scrambles a
107
+ single `00001.ttf`. epubx refused both as "likely Adobe ADEPT", which was wrong
108
+ twice over: the books are perfectly readable, and ADEPT is a different scheme
109
+ that the evidence never named.
110
+
111
+ Distinguish by contents, not presence:
112
+
113
+ | Evidence | Verdict |
114
+ |---|---|
115
+ | `EncryptedKey` present | named DRM — licence-protected |
116
+ | non-font resources encrypted | named DRM, citing the resource |
117
+ | **fonts only** | **not DRM** — parse normally, record `book.obfuscated_fonts` |
118
+
119
+ Do not name a vendor you cannot see. "Likely Adobe ADEPT" was an invented claim.
120
+
121
+ ## 6. Footnote ids live on wrappers
122
+
123
+ ```html
124
+ <li><div id="fn_5" epub:type="footnote"><p>The note text…</p></div></li>
125
+ ```
126
+
127
+ No emission path registers that `<div>`: it is a wrapper, not a block. **1,077
128
+ endnote references in one real book resolved to nothing.** Ownership must be
129
+ resolved from the DOM after the walk, searching, most specific first:
130
+
131
+ 1. the element itself, if a block was emitted from it;
132
+ 2. a block emitted from one of its **descendants** (a `<div>` wrapping a `<p>`);
133
+ 3. the nearest **ancestor-or-self** that emitted a block.
134
+
135
+ Getting this wrong is silent and expensive. Three wrong models were tried
136
+ before the right one:
137
+
138
+ - *descendants first* → the `<section>` wrapper swallows every id; 1,077 notes
139
+ collapse onto ~200 shared blocks;
140
+ - *a "deepest block wins" depth heuristic* → guesswork about which block looks
141
+ nested, still wrong;
142
+ - *keying blocks by `id(element)`* → **lxml recycles element proxies, so `id()`
143
+ aliases.** A `<dd>` resolved to an unrelated `<script>` paragraph. Hold
144
+ elements strongly.
145
+
146
+ A bare `#fn_5` names no document, so cross-document resolution needs an
147
+ id→document index built from raw bytes (ids are *not* unique across a book:
148
+ `#fn_1` recurs in every chapter).
149
+
150
+ ## 7. Guide covers can point at documents
151
+
152
+ `<reference type="cover" href="titlepage.xhtml"/>` names an **XHTML page**, not
153
+ an image. Preferring the guide over `<meta name="cover" content="cover"/>` made
154
+ epubx return `Image(path='titlepage.xhtml', media_type='application/xhtml+xml')`
155
+ — an "image" that is a web page — while ignoring the real `cover.jpeg`.
156
+
157
+ Cover candidates are ordered by authority but must **be images** to win. Three
158
+ conventions coexist: `properties="cover-image"`, `guide`, `<meta name="cover">`.
159
+
160
+ ## 8. Element identity is not stable
161
+
162
+ Covered in §6, but it generalises: **do not key anything on `id(element)`** in
163
+ an lxml tree. Proxies are created and released as you iterate, and a recycled
164
+ address silently attributes data to the wrong node. Key on the element, or hold
165
+ a strong reference.
166
+
167
+ ## 9. Sampling will condemn a readable book
168
+
169
+ A stride sampler (`chapters[::len//8]`) reported a 1,690-chapter C++ textbook
170
+ as image-only. Only **51** chapters held prose; the other 1,639 were page
171
+ images. Every stride landed on an image.
172
+
173
+ Books interleave text and images unpredictably. Walk chapters **in order** and
174
+ stop at the first real prose; then a text book costs only what it takes to find
175
+ its first paragraph.
176
+
177
+ The same book was nearly condemned a second way: its "2.1M characters" were
178
+ **image alt text**, not prose. Count `block.text`, never `plain_text` — that
179
+ folds alt text in, and a scan's alt text can be a whole title page.
180
+
181
+ ## 10. Silent truncation looks like corruption
182
+
183
+ A book opened as `BadZipFile: File is not a zip file`. The archive was fine —
184
+ a concurrent process had truncated the copy mid-write, at exactly 10,485,760
185
+ bytes (10 MB). **Verify the input before blaming the library.** Three of the
186
+ four "corrupt books" investigated in this project were artefacts of the test
187
+ harness, not the corpus.
188
+
189
+ ## 11. Network latency is not parse time
190
+
191
+ `open()` appeared to take **386 ms** over SMB and was about to be "optimised".
192
+ On a local copy the same book opens in **1.2 ms**; the 176 MB book in 13 ms.
193
+
194
+ Benchmark against a local copy, and take a **median over repeats**. A single
195
+ timing over a network mount measures the network.
196
+
197
+ ## 12. Truth-testing an lxml element
198
+
199
+ ```python
200
+ return parse_html(data) or parse_xml(data) # FutureWarning, and wrong on <html/>
201
+ ```
202
+
203
+ An element with no children is falsy. `parse_html` returning a valid empty
204
+ document silently falls through to the XML parser. Test `is not None`.
205
+
206
+ ## 13. Nested blocks must not rewind the ordinal counter
207
+
208
+ Building nested blocks into a temporary list, then restoring the counter along
209
+ with it, hands the same id to a nested block and a later top-level one. Ids
210
+ must come from one monotonic counter per document so `c0000/b0012` is unique
211
+ across the whole chapter — that uniqueness is what makes the id safe to anchor
212
+ annotations and reading positions.
213
+
214
+ ## 14. Zero-length text is not zero content — and not a bug either
215
+
216
+ An "empty" chapter is ambiguous in both directions:
217
+
218
+ - **It may be correct.** In the three *Dr. Slump* manga volumes a typical chapter
219
+ is a single full-page scan: 2 blocks, 1 image, ~3 characters. That is the
220
+ whole page. Nothing is wrong.
221
+ - **It may be a lost-text bug** — and it usually is. During this project, 79
222
+ chapters and then ~130 books reported near-zero text, and *every one* was a
223
+ parser failure (§1), not an empty book.
224
+
225
+ So: before treating an empty chapter as a bug, check whether it holds images or
226
+ is a plausible page; and before treating a book as text-bearing, remember that
227
+ alt attributes count as characters (see §9). "This book has text" deserves to be
228
+ an assertion in a test, not an impression.
229
+
230
+ ## 15. Percent-encoding and `../` in hrefs
231
+
232
+ `src="../OEBPS/img/plate%20one.png"` must normalise to `OEBPS/img/plate one.png`.
233
+ Normalise against the **referring document's** directory, not the OPF's, and
234
+ percent-decode once. Absolute URLs are not zip members and must be skipped —
235
+ silently, since a book referencing a CDN image is still a readable book.
236
+
237
+ ## 16. A walked container drops the text that owns no block
238
+
239
+ The §1 fix handles a `<div>` that holds *only* loose prose. A `<div>` that
240
+ holds loose prose **and** a real block is a different animal, and it cost real
241
+ books their text twice over.
242
+
243
+ `_holds_bare_text` answers "is this wrapper prose, or structure?". A block
244
+ child makes the answer "structure" — correctly, since the `<ul>` must stay a
245
+ list — and the wrapper is then *walked*. But the walk visited child elements
246
+ only, and two kinds of text live in no element at all:
247
+
248
+ - the container's own `.text`, before its first child;
249
+ - every child's `.tail`.
250
+
251
+ *On China* keeps a section's lead-in prose in the div's own text beside a
252
+ nested `<div>`: **67,104 characters dropped, 6% of the book**, while Calibre
253
+ shows every word. A chapter of *Sheng Si Suo* keeps its paragraphs in the
254
+ tails of `<br/>` elements inside a div that also holds a `<ul>`: **13,756
255
+ characters in, 185 out** — and that book's other eight chapters were fine,
256
+ which is exactly why it went unnoticed for so long. Both are §1's rule in a
257
+ new disguise: a block walker that finds no block has not found no content.
258
+
259
+ The fix: `walk` emits the container's own text and every child's tail as
260
+ paragraphs — unless that child's tail is already inside the child's own
261
+ block, which `text_of` arranges by reading an element's tail along with its
262
+ content. `_text()` records that consumption; without the record every
263
+ paragraph's trailing text would be emitted a second time.
264
+
265
+ Measured after: the four chapters recover 13,811 / 5,713 / 6,348 / 10,385
266
+ characters, *On China*'s gap goes to −1%, and the corpus gains 128,052
267
+ characters across 1,663 new paragraphs with **no other block kind changed by
268
+ one**. No fixture held the shape; the audit found it, and Calibre settled it.
269
+
270
+ ---
271
+
272
+ ## 17. Resolving cross-document footnotes can re-enter the parser
273
+
274
+ Resolving a marker whose note lives in another chapter walks that chapter's
275
+ blocks — which parses it. Once compact-id recognition widened the graph
276
+ (`fn674`, `_ftn5` — markers the old classifier never saw), two chapters whose
277
+ notes reference each other re-entered each other's parse: A resolves → walks
278
+ B → B resolves → walks A — and A's `blocks` cached_property was still
279
+ mid-computation, so it *re-ran* instead of returning. Unbounded recursion.
280
+ Four corpus books (Simon & Schuster-style exports) hit `RecursionError` the
281
+ day the recognition shipped; the corpus suite caught all four in one run.
282
+
283
+ Fix: `parse_document` stashes the built (pre-resolution) blocks keyed by
284
+ book + chapter before resolution runs, and `Chapter.blocks` serves the stash
285
+ to reentrant access. Resolution may nest, but every chapter parses exactly
286
+ once, however the references weave. The stash is dropped when the parse
287
+ completes, so nothing lingers.
288
+
289
+ Lesson: widening a classifier widens the *graph* — the resolution order must
290
+ be reentrancy-proof before the widening ships, not after. The corpus guard
291
+ added alongside the recognition is what turned four crashes into one
292
+ afternoon's fix.
293
+
294
+ ---
295
+
296
+ ## 18. `Block.attributes` was a mutable `dict` on a frozen dataclass
297
+
298
+ **Provenance: found by review, not by the corpus.** No book produced this one —
299
+ nothing in the corpus writes to the parsed graph. It is recorded because it is
300
+ the same failure *class* as everything else here: silent, plausible, and a
301
+ disagreement between the graph and the document that nothing reports.
302
+
303
+ `Block` is `@dataclass(frozen=True)`, which stops a field being rebound:
304
+ `block.text = "x"` raises `FrozenInstanceError`. It does **not** stop the dict
305
+ that field points at. `block.attributes["element"] = "p"` succeeded — and
306
+ `ch.blocks` is a memoised `cached_property` whose `Block` objects are also
307
+ handed to the reentrant-parse stash (§17), so one consumer's write was every
308
+ later reader's read, including footnote resolution mid-parse. The graph then
309
+ disagrees with the document, and nothing raises.
310
+
311
+ Fix: `parse_document` returns copies whose `attributes` are wrapped in
312
+ `types.MappingProxyType`, recursing through `items` and `rows`. Freezing runs
313
+ *after* `_resolve_footnotes`, so `target_id` and `dom_ids` are already written.
314
+ Writing to a parsed block's attributes now raises `TypeError`.
315
+
316
+ Two things the fix leans on, both worth keeping in mind:
317
+
318
+ - `parse_stash` deliberately holds the **unfrozen** blocks, because resolution
319
+ still writes to them. `Chapter.blocks` serves those to a reentrant read, and
320
+ the *outer* `cached_property` overwrites the cached value with the frozen
321
+ tuple once its parse returns. That ordering is what makes the freeze survive
322
+ §17's reentrancy, so a test asserts it rather than assuming it.
323
+ - The freeze is a copy, not a conversion. Blocks are built as plain dicts and
324
+ frozen once, at the end of the parse; nothing mutates them afterwards.
325
+
326
+ Lesson: `@dataclass(frozen=True)` is shallow. It stops rebinding, not mutation
327
+ of what a field points at. An immutable public model needs its nested values
328
+ made immutable too.
329
+
330
+ ---
331
+
332
+ ## Known limitations (not bugs)
333
+
334
+ - **Vertical CJK layout** is named as deferred in SPEC.md and is **not
335
+ implemented**. Note what the corpus does and does not say about it: all 300
336
+ books are Chinese or English, but **none declares `writing-mode`**, and their
337
+ stylesheets are plainly horizontal (`text-align: justify`, left/right margins).
338
+ So this corpus provides *no evidence* about vertical layout either way — it is
339
+ untested because the corpus is silent, not because the corpus exercises it.
340
+ An earlier draft of this file claimed the corpus was "almost entirely
341
+ vertical-writing Chinese"; that was an assumption, not a measurement, and it
342
+ was wrong.
343
+ - **Math** has 0 occurrences across the corpus. The MathML path is
344
+ specification-correct and empirically unvalidated — synthetic fixtures only.
345
+ - **5 of 7,905 footnote references cannot resolve.** 2 are in *Zhe Ben Shu Jiao
346
+ Shi Yao*, where the book references `#fn__1`/`#fn__2` but defines
347
+ `#fnt__1`/`#fnt__2` — a publisher typo, missing `t`. The other 3 point at
348
+ external web URLs, which are not zip members. Nothing a parser can do;
349
+ correctly left unresolved rather than guessed at.
350
+ - **Image-only detection is a heuristic** — first-200-chapters-or-first-prose.
351
+ It is deliberately biased toward *not* flagging, because refusing a readable
352
+ book is worse than missing a rare scan.
353
+
354
+ ## How to avoid adding to this list
355
+
356
+ 1. **Run against real books before believing any fix.** Four of these bugs were
357
+ introduced *by* a fix and only caught by re-running the corpus.
358
+ 2. **When a statistic surprises you, read the raw markup** before theorising.
359
+ Twice the surprising statistic was correct and my reading of it was not.
360
+ 3. **Assert absence loudly in tests.** "This book has text" is a test. So is
361
+ "these two notes resolve to different blocks."
362
+ 4. **Prefer naming to guessing.** When the evidence cannot identify a vendor or
363
+ a cause, say what was observed.