epub-extended 0.1.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- epub_extended-0.1.0/LICENSE +21 -0
- epub_extended-0.1.0/MANIFEST.in +12 -0
- epub_extended-0.1.0/PITFALLS.md +363 -0
- epub_extended-0.1.0/PKG-INFO +289 -0
- epub_extended-0.1.0/README.md +261 -0
- epub_extended-0.1.0/RELEASING.md +113 -0
- epub_extended-0.1.0/SPEC.md +276 -0
- epub_extended-0.1.0/epub_extended.egg-info/PKG-INFO +289 -0
- epub_extended-0.1.0/epub_extended.egg-info/SOURCES.txt +29 -0
- epub_extended-0.1.0/epub_extended.egg-info/dependency_links.txt +1 -0
- epub_extended-0.1.0/epub_extended.egg-info/requires.txt +4 -0
- epub_extended-0.1.0/epub_extended.egg-info/top_level.txt +1 -0
- epub_extended-0.1.0/epubx/__init__.py +38 -0
- epub_extended-0.1.0/epubx/content.py +799 -0
- epub_extended-0.1.0/epubx/hrefs.py +36 -0
- epub_extended-0.1.0/epubx/model.py +263 -0
- epub_extended-0.1.0/epubx/nav.py +156 -0
- epub_extended-0.1.0/epubx/package.py +605 -0
- epub_extended-0.1.0/epubx/xmlutil.py +141 -0
- epub_extended-0.1.0/pyproject.toml +43 -0
- epub_extended-0.1.0/setup.cfg +4 -0
- epub_extended-0.1.0/tests/fixtures.py +422 -0
- epub_extended-0.1.0/tests/test_calibre_divs.py +72 -0
- epub_extended-0.1.0/tests/test_corpus.py +112 -0
- epub_extended-0.1.0/tests/test_encoding.py +67 -0
- epub_extended-0.1.0/tests/test_epubx.py +539 -0
- epub_extended-0.1.0/tests/test_loose_text.py +105 -0
- epub_extended-0.1.0/tests/test_serve.py +137 -0
- epub_extended-0.1.0/tests/test_serving.py +100 -0
- epub_extended-0.1.0/tools/audit_corpus.py +422 -0
- epub_extended-0.1.0/tools/serve.py +419 -0
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
MIT License
|
|
2
|
+
|
|
3
|
+
Copyright (c) 2026 Taylor Ren
|
|
4
|
+
|
|
5
|
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
|
6
|
+
of this software and associated documentation files (the "Software"), to deal
|
|
7
|
+
in the Software without restriction, including without limitation the rights
|
|
8
|
+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
|
9
|
+
copies of the Software, and to permit persons to whom the Software is
|
|
10
|
+
furnished to do so, subject to the following conditions:
|
|
11
|
+
|
|
12
|
+
The above copyright notice and this permission notice shall be included in all
|
|
13
|
+
copies or substantial portions of the Software.
|
|
14
|
+
|
|
15
|
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
|
16
|
+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
|
17
|
+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
|
18
|
+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
|
19
|
+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
|
20
|
+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
|
21
|
+
SOFTWARE.
|
|
@@ -0,0 +1,12 @@
|
|
|
1
|
+
# The wheel ships only the package (see [tool.setuptools] in pyproject.toml).
|
|
2
|
+
# The sdist is the source of the project, so it carries the tests, the dev
|
|
3
|
+
# tools, and the docs the README points at.
|
|
4
|
+
include README.md
|
|
5
|
+
include SPEC.md
|
|
6
|
+
include PITFALLS.md
|
|
7
|
+
include RELEASING.md
|
|
8
|
+
|
|
9
|
+
recursive-include tests *.py
|
|
10
|
+
recursive-include tools *.py
|
|
11
|
+
|
|
12
|
+
global-exclude __pycache__ *.py[cod]
|
|
@@ -0,0 +1,363 @@
|
|
|
1
|
+
# epubx — Pitfalls
|
|
2
|
+
|
|
3
|
+
Every entry here is a bug that **epubx actually shipped**, found by running
|
|
4
|
+
against real books. Each one produced plausible-looking output — no exception,
|
|
5
|
+
no warning — and silently lost or misreported content. Synthetic fixtures did
|
|
6
|
+
not catch any of them.
|
|
7
|
+
|
|
8
|
+
Read this before changing the parser. Most of these are ways to be *quietly*
|
|
9
|
+
wrong, which is the only failure mode that matters in a library whose job is to
|
|
10
|
+
report what a document contains.
|
|
11
|
+
|
|
12
|
+
**Evidence base:** 300 EPUB files (1.39 GB, mostly Chinese, from a Calibre
|
|
13
|
+
library). Current state: **0 uncaught exceptions**, 661,108 paragraphs,
|
|
14
|
+
7,900/7,905 footnote references resolved, 4 books named image-only.
|
|
15
|
+
|
|
16
|
+
---
|
|
17
|
+
|
|
18
|
+
## The rule that explains all of them
|
|
19
|
+
|
|
20
|
+
> **A block walker that finds no block has not found no content.**
|
|
21
|
+
|
|
22
|
+
Recursing into a wrapper and emitting nothing looks identical to a wrapper that
|
|
23
|
+
genuinely is empty. Every data-loss bug below is that confusion. When a
|
|
24
|
+
container yields zero blocks, ask what it *actually* held before concluding it
|
|
25
|
+
was empty — the text is usually still there, sitting somewhere you did not look.
|
|
26
|
+
|
|
27
|
+
Corollary, learned the hard way twice: **never infer a book's nature from an
|
|
28
|
+
aggregate.** Sampled chapters, summed character counts and averaged ratios each
|
|
29
|
+
produced a confident, wrong verdict on a real book. Look at one document's raw
|
|
30
|
+
markup before believing any statistic.
|
|
31
|
+
|
|
32
|
+
---
|
|
33
|
+
|
|
34
|
+
## 1. Paragraphs that are not `<p>`
|
|
35
|
+
|
|
36
|
+
**The single largest source of lost text.**
|
|
37
|
+
|
|
38
|
+
`<p>` is not a requirement of XHTML-in-practice. Real books emit paragraphs as:
|
|
39
|
+
|
|
40
|
+
```html
|
|
41
|
+
<!-- Calibre convention: a class, and no <p> anywhere in the file -->
|
|
42
|
+
<div class="p-indent"><span>Body text…</span></div>
|
|
43
|
+
```
|
|
44
|
+
|
|
45
|
+
```html
|
|
46
|
+
<!-- Chinese novels: prose set loose, separated by <br/> -->
|
|
47
|
+
<div class="calibre1">
|
|
48
|
+
<h3>第九部:神秘敵人</h3><br/>
|
|
49
|
+
<br/>
|
|
50
|
+
黃俊和兩個大漢,跟在我們背後…<br/>
|
|
51
|
+
<br/>
|
|
52
|
+
「死神」?不可能的…
|
|
53
|
+
</div>
|
|
54
|
+
```
|
|
55
|
+
|
|
56
|
+
Two separate traps:
|
|
57
|
+
|
|
58
|
+
- **The Calibre class.** Treat every `<div>` as structure and *Killing Lincoln*
|
|
59
|
+
(79 chapters, 511,617 characters) parsed to **zero**. Honour
|
|
60
|
+
`p-indent`/`p-br`/`p-blanc`/`p-continuance`.
|
|
61
|
+
- **Prose in the tails.** The second form's text is in the **`.tail` of each
|
|
62
|
+
`<br/>`**, not in the div's own `.text` (which is indentation whitespace).
|
|
63
|
+
Checking only `.text` finds nothing. ~130 readable books parsed to ~465
|
|
64
|
+
characters each; they now yield their full text.
|
|
65
|
+
|
|
66
|
+
**And the trap in the fix.** The obvious generalisation — "a div with no block
|
|
67
|
+
children is a paragraph" — is wrong: it swallows `<body>` and `<section>` and
|
|
68
|
+
collapses an entire chapter into one block (it broke 10 tests). And "contains a
|
|
69
|
+
block child" is *also* wrong, because the wrapper above legitimately contains an
|
|
70
|
+
`<h3>`. The rule that holds: a block **containing the prose** disqualifies;
|
|
71
|
+
inline markup and headings alongside loose prose do not.
|
|
72
|
+
|
|
73
|
+
## 2. lxml's HTML parser reads undeclared files as latin-1
|
|
74
|
+
|
|
75
|
+
A document with **no charset declaration at all** — valid UTF-8 — came out as
|
|
76
|
+
`Whatâ\x80\x99s next`. lxml sniffs, fails to find a declaration, falls back to
|
|
77
|
+
latin-1, and the damage is invisible until you read the text.
|
|
78
|
+
|
|
79
|
+
**The fix that does not work:** decoding the bytes yourself and re-encoding as
|
|
80
|
+
UTF-8. lxml re-sniffs and guesses latin-1 again. You must pass the encoding to
|
|
81
|
+
the parser:
|
|
82
|
+
|
|
83
|
+
```python
|
|
84
|
+
html.HTMLParser(recover=True, encoding="utf-8") # this is the load-bearing part
|
|
85
|
+
```
|
|
86
|
+
|
|
87
|
+
## 3. `xlink:href` is not namespaced by the HTML parser
|
|
88
|
+
|
|
89
|
+
An SVG-wrapped cover reported **zero images**. The XML parser expands
|
|
90
|
+
`xlink:href` to `{http://www.w3.org/1999/xlink}href`; the **HTML parser leaves
|
|
91
|
+
the literal string `"xlink:href"`**. Reading only the Clark-notation form finds
|
|
92
|
+
nothing, with no error. Accept both.
|
|
93
|
+
|
|
94
|
+
Corollary: books wrap covers in `<svg><image>` rather than `<img>`. Handle SVG
|
|
95
|
+
images or whole pages vanish.
|
|
96
|
+
|
|
97
|
+
## 4. `opf:file-as` is namespaced
|
|
98
|
+
|
|
99
|
+
`<dc:creator opf:file-as="Doe, Jane">` parses to
|
|
100
|
+
`{http://www.idpf.org/2007/opf}file-as`. `element.get("file-as")` returns `None`,
|
|
101
|
+
silently, for every creator in the book.
|
|
102
|
+
|
|
103
|
+
## 5. `encryption.xml` is not DRM
|
|
104
|
+
|
|
105
|
+
Font obfuscation XOR-encodes `.ttf` files to discourage extraction and **leaves
|
|
106
|
+
the text in the clear**. Two corpus books do exactly this — one scrambles a
|
|
107
|
+
single `00001.ttf`. epubx refused both as "likely Adobe ADEPT", which was wrong
|
|
108
|
+
twice over: the books are perfectly readable, and ADEPT is a different scheme
|
|
109
|
+
that the evidence never named.
|
|
110
|
+
|
|
111
|
+
Distinguish by contents, not presence:
|
|
112
|
+
|
|
113
|
+
| Evidence | Verdict |
|
|
114
|
+
|---|---|
|
|
115
|
+
| `EncryptedKey` present | named DRM — licence-protected |
|
|
116
|
+
| non-font resources encrypted | named DRM, citing the resource |
|
|
117
|
+
| **fonts only** | **not DRM** — parse normally, record `book.obfuscated_fonts` |
|
|
118
|
+
|
|
119
|
+
Do not name a vendor you cannot see. "Likely Adobe ADEPT" was an invented claim.
|
|
120
|
+
|
|
121
|
+
## 6. Footnote ids live on wrappers
|
|
122
|
+
|
|
123
|
+
```html
|
|
124
|
+
<li><div id="fn_5" epub:type="footnote"><p>The note text…</p></div></li>
|
|
125
|
+
```
|
|
126
|
+
|
|
127
|
+
No emission path registers that `<div>`: it is a wrapper, not a block. **1,077
|
|
128
|
+
endnote references in one real book resolved to nothing.** Ownership must be
|
|
129
|
+
resolved from the DOM after the walk, searching, most specific first:
|
|
130
|
+
|
|
131
|
+
1. the element itself, if a block was emitted from it;
|
|
132
|
+
2. a block emitted from one of its **descendants** (a `<div>` wrapping a `<p>`);
|
|
133
|
+
3. the nearest **ancestor-or-self** that emitted a block.
|
|
134
|
+
|
|
135
|
+
Getting this wrong is silent and expensive. Three wrong models were tried
|
|
136
|
+
before the right one:
|
|
137
|
+
|
|
138
|
+
- *descendants first* → the `<section>` wrapper swallows every id; 1,077 notes
|
|
139
|
+
collapse onto ~200 shared blocks;
|
|
140
|
+
- *a "deepest block wins" depth heuristic* → guesswork about which block looks
|
|
141
|
+
nested, still wrong;
|
|
142
|
+
- *keying blocks by `id(element)`* → **lxml recycles element proxies, so `id()`
|
|
143
|
+
aliases.** A `<dd>` resolved to an unrelated `<script>` paragraph. Hold
|
|
144
|
+
elements strongly.
|
|
145
|
+
|
|
146
|
+
A bare `#fn_5` names no document, so cross-document resolution needs an
|
|
147
|
+
id→document index built from raw bytes (ids are *not* unique across a book:
|
|
148
|
+
`#fn_1` recurs in every chapter).
|
|
149
|
+
|
|
150
|
+
## 7. Guide covers can point at documents
|
|
151
|
+
|
|
152
|
+
`<reference type="cover" href="titlepage.xhtml"/>` names an **XHTML page**, not
|
|
153
|
+
an image. Preferring the guide over `<meta name="cover" content="cover"/>` made
|
|
154
|
+
epubx return `Image(path='titlepage.xhtml', media_type='application/xhtml+xml')`
|
|
155
|
+
— an "image" that is a web page — while ignoring the real `cover.jpeg`.
|
|
156
|
+
|
|
157
|
+
Cover candidates are ordered by authority but must **be images** to win. Three
|
|
158
|
+
conventions coexist: `properties="cover-image"`, `guide`, `<meta name="cover">`.
|
|
159
|
+
|
|
160
|
+
## 8. Element identity is not stable
|
|
161
|
+
|
|
162
|
+
Covered in §6, but it generalises: **do not key anything on `id(element)`** in
|
|
163
|
+
an lxml tree. Proxies are created and released as you iterate, and a recycled
|
|
164
|
+
address silently attributes data to the wrong node. Key on the element, or hold
|
|
165
|
+
a strong reference.
|
|
166
|
+
|
|
167
|
+
## 9. Sampling will condemn a readable book
|
|
168
|
+
|
|
169
|
+
A stride sampler (`chapters[::len//8]`) reported a 1,690-chapter C++ textbook
|
|
170
|
+
as image-only. Only **51** chapters held prose; the other 1,639 were page
|
|
171
|
+
images. Every stride landed on an image.
|
|
172
|
+
|
|
173
|
+
Books interleave text and images unpredictably. Walk chapters **in order** and
|
|
174
|
+
stop at the first real prose; then a text book costs only what it takes to find
|
|
175
|
+
its first paragraph.
|
|
176
|
+
|
|
177
|
+
The same book was nearly condemned a second way: its "2.1M characters" were
|
|
178
|
+
**image alt text**, not prose. Count `block.text`, never `plain_text` — that
|
|
179
|
+
folds alt text in, and a scan's alt text can be a whole title page.
|
|
180
|
+
|
|
181
|
+
## 10. Silent truncation looks like corruption
|
|
182
|
+
|
|
183
|
+
A book opened as `BadZipFile: File is not a zip file`. The archive was fine —
|
|
184
|
+
a concurrent process had truncated the copy mid-write, at exactly 10,485,760
|
|
185
|
+
bytes (10 MB). **Verify the input before blaming the library.** Three of the
|
|
186
|
+
four "corrupt books" investigated in this project were artefacts of the test
|
|
187
|
+
harness, not the corpus.
|
|
188
|
+
|
|
189
|
+
## 11. Network latency is not parse time
|
|
190
|
+
|
|
191
|
+
`open()` appeared to take **386 ms** over SMB and was about to be "optimised".
|
|
192
|
+
On a local copy the same book opens in **1.2 ms**; the 176 MB book in 13 ms.
|
|
193
|
+
|
|
194
|
+
Benchmark against a local copy, and take a **median over repeats**. A single
|
|
195
|
+
timing over a network mount measures the network.
|
|
196
|
+
|
|
197
|
+
## 12. Truth-testing an lxml element
|
|
198
|
+
|
|
199
|
+
```python
|
|
200
|
+
return parse_html(data) or parse_xml(data) # FutureWarning, and wrong on <html/>
|
|
201
|
+
```
|
|
202
|
+
|
|
203
|
+
An element with no children is falsy. `parse_html` returning a valid empty
|
|
204
|
+
document silently falls through to the XML parser. Test `is not None`.
|
|
205
|
+
|
|
206
|
+
## 13. Nested blocks must not rewind the ordinal counter
|
|
207
|
+
|
|
208
|
+
Building nested blocks into a temporary list, then restoring the counter along
|
|
209
|
+
with it, hands the same id to a nested block and a later top-level one. Ids
|
|
210
|
+
must come from one monotonic counter per document so `c0000/b0012` is unique
|
|
211
|
+
across the whole chapter — that uniqueness is what makes the id safe to anchor
|
|
212
|
+
annotations and reading positions.
|
|
213
|
+
|
|
214
|
+
## 14. Zero-length text is not zero content — and not a bug either
|
|
215
|
+
|
|
216
|
+
An "empty" chapter is ambiguous in both directions:
|
|
217
|
+
|
|
218
|
+
- **It may be correct.** In the three *Dr. Slump* manga volumes a typical chapter
|
|
219
|
+
is a single full-page scan: 2 blocks, 1 image, ~3 characters. That is the
|
|
220
|
+
whole page. Nothing is wrong.
|
|
221
|
+
- **It may be a lost-text bug** — and it usually is. During this project, 79
|
|
222
|
+
chapters and then ~130 books reported near-zero text, and *every one* was a
|
|
223
|
+
parser failure (§1), not an empty book.
|
|
224
|
+
|
|
225
|
+
So: before treating an empty chapter as a bug, check whether it holds images or
|
|
226
|
+
is a plausible page; and before treating a book as text-bearing, remember that
|
|
227
|
+
alt attributes count as characters (see §9). "This book has text" deserves to be
|
|
228
|
+
an assertion in a test, not an impression.
|
|
229
|
+
|
|
230
|
+
## 15. Percent-encoding and `../` in hrefs
|
|
231
|
+
|
|
232
|
+
`src="../OEBPS/img/plate%20one.png"` must normalise to `OEBPS/img/plate one.png`.
|
|
233
|
+
Normalise against the **referring document's** directory, not the OPF's, and
|
|
234
|
+
percent-decode once. Absolute URLs are not zip members and must be skipped —
|
|
235
|
+
silently, since a book referencing a CDN image is still a readable book.
|
|
236
|
+
|
|
237
|
+
## 16. A walked container drops the text that owns no block
|
|
238
|
+
|
|
239
|
+
The §1 fix handles a `<div>` that holds *only* loose prose. A `<div>` that
|
|
240
|
+
holds loose prose **and** a real block is a different animal, and it cost real
|
|
241
|
+
books their text twice over.
|
|
242
|
+
|
|
243
|
+
`_holds_bare_text` answers "is this wrapper prose, or structure?". A block
|
|
244
|
+
child makes the answer "structure" — correctly, since the `<ul>` must stay a
|
|
245
|
+
list — and the wrapper is then *walked*. But the walk visited child elements
|
|
246
|
+
only, and two kinds of text live in no element at all:
|
|
247
|
+
|
|
248
|
+
- the container's own `.text`, before its first child;
|
|
249
|
+
- every child's `.tail`.
|
|
250
|
+
|
|
251
|
+
*On China* keeps a section's lead-in prose in the div's own text beside a
|
|
252
|
+
nested `<div>`: **67,104 characters dropped, 6% of the book**, while Calibre
|
|
253
|
+
shows every word. A chapter of *Sheng Si Suo* keeps its paragraphs in the
|
|
254
|
+
tails of `<br/>` elements inside a div that also holds a `<ul>`: **13,756
|
|
255
|
+
characters in, 185 out** — and that book's other eight chapters were fine,
|
|
256
|
+
which is exactly why it went unnoticed for so long. Both are §1's rule in a
|
|
257
|
+
new disguise: a block walker that finds no block has not found no content.
|
|
258
|
+
|
|
259
|
+
The fix: `walk` emits the container's own text and every child's tail as
|
|
260
|
+
paragraphs — unless that child's tail is already inside the child's own
|
|
261
|
+
block, which `text_of` arranges by reading an element's tail along with its
|
|
262
|
+
content. `_text()` records that consumption; without the record every
|
|
263
|
+
paragraph's trailing text would be emitted a second time.
|
|
264
|
+
|
|
265
|
+
Measured after: the four chapters recover 13,811 / 5,713 / 6,348 / 10,385
|
|
266
|
+
characters, *On China*'s gap goes to −1%, and the corpus gains 128,052
|
|
267
|
+
characters across 1,663 new paragraphs with **no other block kind changed by
|
|
268
|
+
one**. No fixture held the shape; the audit found it, and Calibre settled it.
|
|
269
|
+
|
|
270
|
+
---
|
|
271
|
+
|
|
272
|
+
## 17. Resolving cross-document footnotes can re-enter the parser
|
|
273
|
+
|
|
274
|
+
Resolving a marker whose note lives in another chapter walks that chapter's
|
|
275
|
+
blocks — which parses it. Once compact-id recognition widened the graph
|
|
276
|
+
(`fn674`, `_ftn5` — markers the old classifier never saw), two chapters whose
|
|
277
|
+
notes reference each other re-entered each other's parse: A resolves → walks
|
|
278
|
+
B → B resolves → walks A — and A's `blocks` cached_property was still
|
|
279
|
+
mid-computation, so it *re-ran* instead of returning. Unbounded recursion.
|
|
280
|
+
Four corpus books (Simon & Schuster-style exports) hit `RecursionError` the
|
|
281
|
+
day the recognition shipped; the corpus suite caught all four in one run.
|
|
282
|
+
|
|
283
|
+
Fix: `parse_document` stashes the built (pre-resolution) blocks keyed by
|
|
284
|
+
book + chapter before resolution runs, and `Chapter.blocks` serves the stash
|
|
285
|
+
to reentrant access. Resolution may nest, but every chapter parses exactly
|
|
286
|
+
once, however the references weave. The stash is dropped when the parse
|
|
287
|
+
completes, so nothing lingers.
|
|
288
|
+
|
|
289
|
+
Lesson: widening a classifier widens the *graph* — the resolution order must
|
|
290
|
+
be reentrancy-proof before the widening ships, not after. The corpus guard
|
|
291
|
+
added alongside the recognition is what turned four crashes into one
|
|
292
|
+
afternoon's fix.
|
|
293
|
+
|
|
294
|
+
---
|
|
295
|
+
|
|
296
|
+
## 18. `Block.attributes` was a mutable `dict` on a frozen dataclass
|
|
297
|
+
|
|
298
|
+
**Provenance: found by review, not by the corpus.** No book produced this one —
|
|
299
|
+
nothing in the corpus writes to the parsed graph. It is recorded because it is
|
|
300
|
+
the same failure *class* as everything else here: silent, plausible, and a
|
|
301
|
+
disagreement between the graph and the document that nothing reports.
|
|
302
|
+
|
|
303
|
+
`Block` is `@dataclass(frozen=True)`, which stops a field being rebound:
|
|
304
|
+
`block.text = "x"` raises `FrozenInstanceError`. It does **not** stop the dict
|
|
305
|
+
that field points at. `block.attributes["element"] = "p"` succeeded — and
|
|
306
|
+
`ch.blocks` is a memoised `cached_property` whose `Block` objects are also
|
|
307
|
+
handed to the reentrant-parse stash (§17), so one consumer's write was every
|
|
308
|
+
later reader's read, including footnote resolution mid-parse. The graph then
|
|
309
|
+
disagrees with the document, and nothing raises.
|
|
310
|
+
|
|
311
|
+
Fix: `parse_document` returns copies whose `attributes` are wrapped in
|
|
312
|
+
`types.MappingProxyType`, recursing through `items` and `rows`. Freezing runs
|
|
313
|
+
*after* `_resolve_footnotes`, so `target_id` and `dom_ids` are already written.
|
|
314
|
+
Writing to a parsed block's attributes now raises `TypeError`.
|
|
315
|
+
|
|
316
|
+
Two things the fix leans on, both worth keeping in mind:
|
|
317
|
+
|
|
318
|
+
- `parse_stash` deliberately holds the **unfrozen** blocks, because resolution
|
|
319
|
+
still writes to them. `Chapter.blocks` serves those to a reentrant read, and
|
|
320
|
+
the *outer* `cached_property` overwrites the cached value with the frozen
|
|
321
|
+
tuple once its parse returns. That ordering is what makes the freeze survive
|
|
322
|
+
§17's reentrancy, so a test asserts it rather than assuming it.
|
|
323
|
+
- The freeze is a copy, not a conversion. Blocks are built as plain dicts and
|
|
324
|
+
frozen once, at the end of the parse; nothing mutates them afterwards.
|
|
325
|
+
|
|
326
|
+
Lesson: `@dataclass(frozen=True)` is shallow. It stops rebinding, not mutation
|
|
327
|
+
of what a field points at. An immutable public model needs its nested values
|
|
328
|
+
made immutable too.
|
|
329
|
+
|
|
330
|
+
---
|
|
331
|
+
|
|
332
|
+
## Known limitations (not bugs)
|
|
333
|
+
|
|
334
|
+
- **Vertical CJK layout** is named as deferred in SPEC.md and is **not
|
|
335
|
+
implemented**. Note what the corpus does and does not say about it: all 300
|
|
336
|
+
books are Chinese or English, but **none declares `writing-mode`**, and their
|
|
337
|
+
stylesheets are plainly horizontal (`text-align: justify`, left/right margins).
|
|
338
|
+
So this corpus provides *no evidence* about vertical layout either way — it is
|
|
339
|
+
untested because the corpus is silent, not because the corpus exercises it.
|
|
340
|
+
An earlier draft of this file claimed the corpus was "almost entirely
|
|
341
|
+
vertical-writing Chinese"; that was an assumption, not a measurement, and it
|
|
342
|
+
was wrong.
|
|
343
|
+
- **Math** has 0 occurrences across the corpus. The MathML path is
|
|
344
|
+
specification-correct and empirically unvalidated — synthetic fixtures only.
|
|
345
|
+
- **5 of 7,905 footnote references cannot resolve.** 2 are in *Zhe Ben Shu Jiao
|
|
346
|
+
Shi Yao*, where the book references `#fn__1`/`#fn__2` but defines
|
|
347
|
+
`#fnt__1`/`#fnt__2` — a publisher typo, missing `t`. The other 3 point at
|
|
348
|
+
external web URLs, which are not zip members. Nothing a parser can do;
|
|
349
|
+
correctly left unresolved rather than guessed at.
|
|
350
|
+
- **Image-only detection is a heuristic** — first-200-chapters-or-first-prose.
|
|
351
|
+
It is deliberately biased toward *not* flagging, because refusing a readable
|
|
352
|
+
book is worse than missing a rare scan.
|
|
353
|
+
|
|
354
|
+
## How to avoid adding to this list
|
|
355
|
+
|
|
356
|
+
1. **Run against real books before believing any fix.** Four of these bugs were
|
|
357
|
+
introduced *by* a fix and only caught by re-running the corpus.
|
|
358
|
+
2. **When a statistic surprises you, read the raw markup** before theorising.
|
|
359
|
+
Twice the surprising statistic was correct and my reading of it was not.
|
|
360
|
+
3. **Assert absence loudly in tests.** "This book has text" is a test. So is
|
|
361
|
+
"these two notes resolve to different blocks."
|
|
362
|
+
4. **Prefer naming to guessing.** When the evidence cannot identify a vendor or
|
|
363
|
+
a cause, say what was observed.
|