single-file-core 1.5.119 → 1.5.121
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/core/index.js +4 -0
- package/core/lib/processor-helper-inline.js +3 -3
- package/core/util.js +3 -1
- package/doc/assets/singlefile-archive-byte-map.svg +195 -217
- package/doc/singlefile-archive.md +439 -180
- package/eslint.config.mjs +6 -0
- package/package.json +2 -2
- package/processors/compression/compression-display.js +0 -11
- package/processors/compression/compression-extract.js +0 -4
- package/processors/compression/compression-packager.js +0 -12
- package/processors/compression/compression-router.js +0 -35
- package/processors/compression/compression.js +105 -181
- package/processors/hooks/content/content-hooks-frames-web.js +2 -1
- package/processors/lazy/content/content-lazy-loader.js +21 -18
- package/test/sfz-harness/README.md +10 -1
- package/test/sfz-harness/byte-map.js +137 -0
- package/test/sfz-harness/format-rules.js +132 -2
- package/test/sfz-harness/option-wiring.js +0 -3
- package/test/sfz-harness/zip64.js +77 -0
|
@@ -3,7 +3,7 @@
|
|
|
3
3
|
**Status: draft.** This document specifies the SingleFile archive, the polyglot file
|
|
4
4
|
format produced by [SingleFile](https://github.com/gildas-lormeau/SingleFile) when it
|
|
5
5
|
saves a page as a ZIP archive. It is written against the reference implementation,
|
|
6
|
-
[single-file-core](https://github.com/gildas-lormeau/single-file-core) 1.5.
|
|
6
|
+
[single-file-core](https://github.com/gildas-lormeau/single-file-core) 1.5.120
|
|
7
7
|
(`processors/compression/`), and every byte-level statement has been verified on
|
|
8
8
|
generated specimen files.
|
|
9
9
|
|
|
@@ -62,8 +62,9 @@ little software as possible. Each way of opening the file has a simpler fallback
|
|
|
62
62
|
is hidden by construction, so the browser displays neither the page nor raw
|
|
63
63
|
archive bytes (§4.1).
|
|
64
64
|
3. Renamed to `.zip`, the file opens in a ZIP tool; the page and each resource are
|
|
65
|
-
ordinary entries.
|
|
66
|
-
|
|
65
|
+
ordinary entries. Two measured readers refuse a self-extracting variant even from
|
|
66
|
+
seekable input; the ranking below names them. A forward-only reader refuses it too,
|
|
67
|
+
and is a non-goal (§1.2).
|
|
67
68
|
4. Renamed to `.pdf` or `.png` (when those faces are present), the file opens in a PDF
|
|
68
69
|
viewer or an image viewer.
|
|
69
70
|
|
|
@@ -103,13 +104,15 @@ Three consequences shape everything below:
|
|
|
103
104
|
|
|
104
105
|
### 1.2 Non-goals
|
|
105
106
|
|
|
106
|
-
- **Forward-only ZIP parsers.**
|
|
107
|
-
entries are preceded by non-ZIP bytes
|
|
108
|
-
offset 0
|
|
107
|
+
- **Forward-only ZIP parsers.** Every archive with a face requires central-directory-driven
|
|
108
|
+
reading, since the entries are then preceded by non-ZIP bytes. The variant with no other
|
|
109
|
+
face is an ordinary ZIP file and streams from offset 0 like any other (§8.1, class A);
|
|
110
|
+
parsers that require that are out of scope for the rest (§7).
|
|
109
111
|
- **In-place modification by generic ZIP tools.** The face invariants are global:
|
|
110
112
|
the writer picks each hiding tag only after checking the exact bytes it must hide,
|
|
111
|
-
the recovery payload of universal mode contains a checksum of the
|
|
112
|
-
|
|
113
|
+
the recovery payload of universal mode contains a checksum of the ZIP region
|
|
114
|
+
without its comment-length field, and the PDF and PNG structures wrap the
|
|
115
|
+
archive (§5.4). A tool
|
|
113
116
|
that adds, removes or recompresses entries invalidates them, and most rewriters drop
|
|
114
117
|
the prepended and appended regions outright. A generically rewritten file keeps at
|
|
115
118
|
best its ZIP face. Editing an archive means producing a new one through the writer
|
|
@@ -117,7 +120,8 @@ Three consequences shape everything below:
|
|
|
117
120
|
- **Multi-page archives.** The reference implementation can bundle several saved
|
|
118
121
|
pages into one archive behind a routing bootstrap (`multiPageArchive`). This
|
|
119
122
|
version of the document specifies single-page archives only; the multi-page
|
|
120
|
-
layout is out of scope
|
|
123
|
+
layout is out of scope, and so are the regions it adds to the prologue, which this
|
|
124
|
+
document does not describe.
|
|
121
125
|
- **Confidentiality outside the ZIP entries.** A password encrypts ZIP entry contents
|
|
122
126
|
only (AES). The PDF and PNG faces render the page content and are plaintext by
|
|
123
127
|
design; the writer withholds what it can without breaking a face, as described in
|
|
@@ -133,12 +137,12 @@ Three consequences shape everything below:
|
|
|
133
137
|
| **universal mode** | The variant whose HTML face can extract the archive from the *parsed page text*, the text and comment nodes the HTML parser produced, and therefore needs no access to its own raw bytes. Named "universal" because it works from any location, including the `file:` protocol. |
|
|
134
138
|
| **wrapper tag** | The HTML construct that hides a binary region from the HTML parser, `<!--`…`-->` by default (§5.1). |
|
|
135
139
|
| **appended data** | Bytes after the ZIP End Of Central Directory record. Readers tolerate them within the window their EOCD scan already covers: 65557 bytes from the end of the file (the 22-byte record plus the 65535-byte maximum comment length); "the 64 KB window" refers to this. It may be left undeclared or declared as the archive comment; both forms are valid ZIP and readers MUST accept both (§4.2). The recovery payload can be computed before that choice is made because it stops two bytes short of the record, excluding its comment-length field (see *recovered range* below). |
|
|
136
|
-
| **ZIP region** | The contiguous byte range holding the archive proper: from the first local file header the ZIP writer emitted through the last byte of the End Of Central Directory record. It spans the `zip-entries`, `pdf-central-record` (when present) and `central-directory · eocd` blocks of §3, and in the HTML variants it is
|
|
140
|
+
| **ZIP region** | The contiguous byte range holding the archive proper: from the first local file header the ZIP writer emitted through the last byte of the End Of Central Directory record. It spans the `zip-entries`, `pdf-central-record` (when present) and `central-directory · eocd` blocks of §3, and in the HTML variants it is the content of the last wrapper, exactly so on the element rungs and preceded by the `sfz-data` identifier on the comment rung, which the extractor steps over. It does **not** include `pdf-local-header` or the PDF document, which sit earlier in the file. |
|
|
137
141
|
| **archive** | The *logical* ZIP file: the set of entries the central directory describes, wherever their bytes lie. This is distinct from the ZIP region above, which is a contiguous byte range. Every entry but one has its bytes inside the region; `page.pdf` is the deliberate exception, an entry of the archive whose local header and data sit before the region (§4.2). "Archive" in this document always means the logical file, "ZIP region" always the byte range, and the two differ only in the PDF-with-HTML variants. |
|
|
138
142
|
| **recovered range** | What the universal extractor reproduces (§4.5): the ZIP region minus its last two bytes, the comment-length field of the End Of Central Directory record. That field is the one part of the record whose value depends on what follows the region, so leaving it out is what lets a writer decide the appended-data form after the recovery payload is final (§4.2). The extractor supplies the two bytes itself, as zeroes — the recovered range carries no comment. |
|
|
139
143
|
| **reference writer** | `createArchive()` in single-file-core `processors/compression/compression.js`. |
|
|
140
144
|
| **bootstrap** | The inline script in the HTML face that locates, extracts and displays the archived page. |
|
|
141
|
-
| **`sfz` identifiers** | Three byte-level identifiers
|
|
145
|
+
| **`sfz` identifiers** | Three byte-level identifiers carrying the `sfz` prefix matter to this document: `data-sfz`, `<sfz-extra-data>` and `sfz-data`. A conforming file may hold others the format says nothing about — the reference bootstrap gives its own status messages `sfz`-prefixed ids, which no reader has any reason to look for. The prefix is inherited from SingleFileZ, the browser extension the format originated in (since merged into SingleFile), and is kept unchanged for compatibility with existing files. They are wire identifiers, not the format's name. Two of the three have a role in the format: `<sfz-extra-data>` is the element carrying the recovery payload, and `sfz-data` is the identifier the universal extractor addresses the ZIP region with — an `id` attribute on the wrapper element, or the first characters of the wrapper comment's data (§4.5). `data-sfz` is a marker the reference writer happens to put on the root element; this document mentions it only where it describes bytes those files contain. |
|
|
142
146
|
|
|
143
147
|
## 2. Variants: composing faces
|
|
144
148
|
|
|
@@ -164,11 +168,17 @@ include it because the clients producing those variants enable it by default, bu
|
|
|
164
168
|
`embeddedImage` and `embeddedPdf` compose with a non-universal self-extracting file
|
|
165
169
|
just as well; the extension is then `.zip.html`. The one interaction: `page.pdf` lies
|
|
166
170
|
outside the ZIP region (§1.3), so the page-text extraction path does not recover it
|
|
167
|
-
with the rest
|
|
168
|
-
|
|
171
|
+
with the rest. The extractor therefore filters that entry out unconditionally, on
|
|
172
|
+
every acquisition path including the ones that read raw bytes and could return it
|
|
173
|
+
(§4.5). In the last three rows there is no
|
|
169
174
|
HTML face, so the option does not apply. The *Specimen* column names the measured
|
|
170
175
|
reference files this document cites; §8 records how to regenerate them.
|
|
171
176
|
|
|
177
|
+
Other writer options shape the file without adding a face: `preventAppendedData` and
|
|
178
|
+
`declareAppendedData` (§4.2, §5.2), `includeBOM` (§3.1), `insertTextBody` (§4.6),
|
|
179
|
+
`password` (§5.6), `createRootDirectory` (§7.1), and the head-element switches
|
|
180
|
+
`insertCanonicalLink`, `insertMetaNoIndex` and `insertMetaCSP` (§3.1).
|
|
181
|
+
|
|
172
182
|
Notes on composition:
|
|
173
183
|
|
|
174
184
|
- **Universal mode requires the HTML face** (it is a property of the bootstrap) and is
|
|
@@ -235,8 +245,8 @@ they are out of reach. Taking the three in turn:
|
|
|
235
245
|
(`includeBOM`), where nothing depends on the declared charset.
|
|
236
246
|
- A **user override** is the one no software can prevent, and the rarest.
|
|
237
247
|
|
|
238
|
-
On `file:` URLs the bootstrap goes straight to page-text extraction, since
|
|
239
|
-
|
|
248
|
+
On `file:` URLs the bootstrap goes straight to page-text extraction, since it attempts
|
|
249
|
+
no raw read there (§4.1) — but there is also no transport layer, so the first of the
|
|
240
250
|
three cannot arise on the very path that depends on the charset most.
|
|
241
251
|
|
|
242
252
|
The failure is safe rather than silent, which is why the precondition is worth stating
|
|
@@ -267,12 +277,39 @@ or overridden by a server still tends to be decoded the way the extractor expect
|
|
|
267
277
|
also keeps the reverse table small, at 27 entries (§5.5).
|
|
268
278
|
|
|
269
279
|
The second part is the extra-data payload, which does **not**
|
|
270
|
-
contain the archive: it carries only what the round trip destroys
|
|
271
|
-
|
|
280
|
+
contain the archive: it carries only what the round trip destroys or leaves
|
|
281
|
+
undetermined, namely a checksum, the recovered range's length, and the information
|
|
282
|
+
needed to restore newline bytes, which the parser normalizes.
|
|
272
283
|
The parser also replaces NUL bytes with U+FFFD; since no byte decodes to U+FFFD under
|
|
273
284
|
a qualifying encoding, the extractor maps U+FFFD back to NUL unambiguously and the
|
|
274
285
|
payload needs nothing for it (§5.5).
|
|
275
286
|
|
|
287
|
+
The declared charset governs the **whole document**, not only the regions the format
|
|
288
|
+
reasons about. Everything the parser reads is decoded with it, the bootstrap script
|
|
289
|
+
included, and the writer's own code is therefore subject to the same single-byte
|
|
290
|
+
decoding as the page it carries. In universal mode the bootstrap MUST contain no
|
|
291
|
+
character outside ASCII.
|
|
292
|
+
|
|
293
|
+
Unlike the `<title>` (§4.6), it cannot be rescued by
|
|
294
|
+
character references. A `<script>` element's content is script data, a tokenizer state
|
|
295
|
+
that does not resolve them: `☺` written there stays seven literal characters and
|
|
296
|
+
reaches the program as seven characters. The escape has to happen one level down, in
|
|
297
|
+
the JavaScript source — `\u263A` rather than `☺`, an escape the language resolves when
|
|
298
|
+
the script is compiled, not one the HTML parser resolves when the file is read. A
|
|
299
|
+
minifier will undo this if allowed to, since printing the shortest form is its default
|
|
300
|
+
and the shortest form of `\u263A` is the literal character. A writer that assembles
|
|
301
|
+
the bootstrap through a minifier MUST configure it to emit ASCII only.
|
|
302
|
+
|
|
303
|
+
The consequence of getting this wrong is worse than the mojibake a raw title produces,
|
|
304
|
+
and that is the reason for the MUST. A garbled title is visible; a garbled string
|
|
305
|
+
inside the extractor is not. A lookup table is the sharpest case. Emitted as literal
|
|
306
|
+
characters, a CP437 table is re-decoded as windows-1252 and grows from 256 entries to
|
|
307
|
+
508, shifting every lookup past the first 32 by 60 positions — and nothing about the
|
|
308
|
+
page looks wrong, because the damage is confined to names the table decodes. Under
|
|
309
|
+
§5.8 that can be a single entry, and one is enough when it is the entry the extractor
|
|
310
|
+
matches by name. The requirement is on the whole bootstrap rather than on any table
|
|
311
|
+
inside it, because a minifier does not know which strings are load-bearing.
|
|
312
|
+
|
|
276
313
|
### 2.2 File name conventions
|
|
277
314
|
|
|
278
315
|
The reference implementation names files by variant: `.zip` (no HTML face),
|
|
@@ -281,7 +318,8 @@ conventions for humans and pickers; **readers MUST NOT rely on the file name**.
|
|
|
281
318
|
face is discoverable from the bytes alone: PNG and PDF by their signatures, the ZIP
|
|
282
319
|
face by its End Of Central Directory record, and the HTML face by an `<html` start tag
|
|
283
320
|
occurring before the first local file header — inside the first `tEXt` chunk's data in
|
|
284
|
-
the PNG variants, where the markup begins after the chunk's keyword
|
|
321
|
+
the PNG variants, where the markup begins after the chunk's keyword and its NUL
|
|
322
|
+
separator. Inside the
|
|
285
323
|
archive, the `index.html` and `manifest.json` entries mark it as a saved page; a
|
|
286
324
|
reader should identify it that way (§7.1). The self-extracting variants are told apart
|
|
287
325
|
the same way: only a universal file carries an `<sfz-extra-data>` element.
|
|
@@ -289,9 +327,10 @@ the same way: only a universal file carries an `<sfz-extra-data>` element.
|
|
|
289
327
|
## 3. The byte map
|
|
290
328
|
|
|
291
329
|
Unless a row states otherwise, the layouts below are measured from specimen files
|
|
292
|
-
saved from `example.com` (the generation commands are in §8)
|
|
293
|
-
|
|
294
|
-
|
|
330
|
+
saved from `example.com` (the generation commands are in §8). The relocated row covers
|
|
331
|
+
two cases with one layout, `preventAppendedData` and a payload over 64 KB: the first is
|
|
332
|
+
measured on the relocated specimen, the second derived from the writer rules, because
|
|
333
|
+
such a payload requires an archive too large for a readable specimen. The figure below shows
|
|
295
334
|
the regions and their order; the glossary of §3.1 is the normative list, and it states
|
|
296
335
|
in text everything the figure conveys.
|
|
297
336
|
|
|
@@ -309,21 +348,21 @@ face adds, then the regions the PNG face adds.
|
|
|
309
348
|
|
|
310
349
|
| Region | Producer | Present | Contents |
|
|
311
350
|
|---|---|---|---|
|
|
312
|
-
| `html-prologue` | HTML | HTML face | Doctype, the root element start tag, an optional implementation-defined comment,
|
|
313
|
-
| `bootstrap` | HTML | HTML face | One inline `<script>`: the embedded ZIP reader, the extractor, the display routine, and the content-acquisition logic (§4.1). The wrapper start tag that opens the ZIP region follows it, directly or after a relocated `extra-data`. |
|
|
314
|
-
| `<!--` / `-->` | HTML | HTML face | The wrapper tag pair hiding a binary region from the HTML parser — comment tags by default, another pair when the hidden bytes
|
|
351
|
+
| `html-prologue` | HTML | HTML face | Doctype, the root element start tag, `<meta charset>`, an optional implementation-defined comment, title, optional head elements (canonical link, `robots` meta, viewport, Content-Security-Policy), minimal CSS, `<body hidden>`, wait/error messages, optional text body (§4.6). The leading comment, the title, the canonical link and the text body are withheld when a password is set (§5.6). In the plain variant an optional UTF-8 BOM MAY precede the doctype (`includeBOM`); universal and PNG variants never carry one, and the reference writer ignores the option there. In the PNG variants the region is split: everything through `<body hidden>` is the data of the `tEXt "PNG"` chunk, while the messages and the optional text body follow the `tEXt "ZIP"` chunk header; the doctype and the leading comment are dropped. |
|
|
352
|
+
| `bootstrap` | HTML | HTML face | One inline `<script>`: the embedded ZIP reader, the extractor, the display routine, and the content-acquisition logic (§4.1). In universal mode its bytes MUST be pure ASCII, since the declared charset decodes this region like any other and character references do not apply inside script data (§2.1). The wrapper start tag that opens the ZIP region follows it, directly or after a relocated `extra-data`. |
|
|
353
|
+
| `<!--` / `-->` | HTML | HTML face | The wrapper tag pair hiding a binary region from the HTML parser — comment tags by default, another pair when the hidden bytes defeat them — which `-->` is only the commonest way to do, the full test being `<!--`, `--!>`, a trailing `<!-` and, for the PNG payload, a leading `>` or `->` (§5.1). Drawn at each opening and closing position. The close tag is absent whenever the recovery payload is relocated (§5.2): under `preventAppendedData`, when the payload outgrows the appended-data budget, or on the `<plaintext>` wrapper which cannot close. No markup then follows the archive and the wrapper runs to end-of-file. That does not mean the file ends at the EOCD — the PNG face's tail still follows, inside the wrapper, where it parses as text (§5.1). |
|
|
315
354
|
| `zip-entries` | ZIP | always | The archive's local file headers and entry data, written by the ZIP writer. The central directory of an archive written by the reference writer lists `index.html` (the page) first, then `manifest.json` (a JSON description of the archive: original URL, title, save time, resource-to-URL map — informative; the page displays without it), then the resources; the *physical* order of the local headers inside the region is not guaranteed to match, and readers MUST NOT rely on either order — entries are addressed by name (§7.1). |
|
|
316
355
|
| `central-directory · eocd` | ZIP | always | The central-directory records followed by the End Of Central Directory record. All offsets are absolute file positions (§5.3). In the HTML+PDF variants the EOCD accounts for the injected `pdf-central-record` (how the writer achieves that is §6). |
|
|
317
356
|
| `extra-data` | extractor | universal | `<sfz-extra-data>` element holding the base64, deflate-compressed recovery payload (§5.5). It always sits outside the wrapper, so it parses as a real element the extractor can address. Normal placement: after the EOCD, between the wrapper close tag and the end tags. Relocated placement, used when the payload exceeds the 64 KB appended-data window or `preventAppendedData` is set: immediately before the wrapper start tag. In the relocated form the element is followed by space padding: its room is reserved before the archive is written, because the region precedes the ZIP data and resizing it would shift every central-directory offset (§6). Neither placement carries positional meaning — the extractor finds the ZIP region by identifier, not relative to this element (§4.5). |
|
|
318
|
-
| `</body></html>` | HTML | HTML face | The end tags closing the document after the wrapper close tag. Omitted
|
|
319
|
-
| `pdf-local-header` | ZIP | PDF face with HTML | The hand-built local file header for `page.pdf` (STORE, checksum precomputed), written immediately before the PDF document so ZIP readers see an ordinary entry whose data is the PDF (§6). |
|
|
320
|
-
| `pdf-document` | PDF | PDF face | The raw PDF bytes. With the HTML face, wrapped together with `pdf-local-header` in a wrapper tag pair inside `html-prologue`, placed so `%PDF-` starts at offset 1024 or lower — the range PDF readers search for the header, which is what lets a PDF document start after other bytes at all (§4.3). Without the HTML face the file simply *starts* with the PDF document, as prepended data the ZIP face tolerates; `page.pdf` is then not an archive entry at all — no local header, no central record. |
|
|
357
|
+
| `</body></html>` | HTML | HTML face | The end tags closing the document after the wrapper close tag. Omitted whenever the recovery payload is relocated (§5.2), and in the PNG variants so the file can end with the PNG tail. |
|
|
358
|
+
| `pdf-local-header` | ZIP | PDF face with HTML | The hand-built local file header for `page.pdf` (STORE, checksum precomputed, language encoding flag set as on every other entry — §5.8), written immediately before the PDF document so ZIP readers see an ordinary entry whose data is the PDF (§6). |
|
|
359
|
+
| `pdf-document` | PDF | PDF face | The raw PDF bytes. With the HTML face, wrapped together with `pdf-local-header` in a wrapper tag pair inside `html-prologue`, placed so `%PDF-` starts at offset 1024 or lower — the range PDF readers search for the header, which is what lets a PDF document start after other bytes at all (§4.3). Without the HTML face and without the PNG face, the file simply *starts* with the PDF document, as prepended data the ZIP face tolerates; `page.pdf` is then not an archive entry at all — no local header, no central record. |
|
|
321
360
|
| `pdf-central-record` | ZIP | PDF face with HTML | The central-directory record for `page.pdf`, injected *before* the writer's own central directory. The start of the central directory is the one place a record can be added without moving any offset the writer already committed, and it makes `page.pdf` the first entry ZIP tools list (§6). |
|
|
322
361
|
| `png-signature · IHDR` | PNG | PNG face | The 8-byte PNG signature and the `IHDR` chunk declaring the source image's dimensions — the first 33 bytes of the file. |
|
|
323
|
-
| `tEXt "PNG"` | PNG | PNG face with HTML | The
|
|
324
|
-
| `tEXt "PDF"` | PNG | PNG + PDF faces without HTML | The length, type
|
|
325
|
-
| `pixel-data chunks` | PNG | PNG face |
|
|
326
|
-
| `tEXt "ZIP"` | PNG | PNG face | The length, type
|
|
362
|
+
| `tEXt "PNG"` | PNG | PNG face with HTML | The 12 header bytes of the first `tEXt` chunk: the 4-byte big-endian length, the type, the keyword and its NUL separator. Its data is `html-prologue` (with the PDF face, the embedded PDF document rides inside it too), ending with the wrapper start tag. |
|
|
363
|
+
| `tEXt "PDF"` | PNG | PNG + PDF faces without HTML | The 12 header bytes (length, type, keyword, NUL separator) of a `tEXt` chunk whose data is the raw PDF document. Written only when the PNG and PDF faces combine without HTML — with the HTML face the PDF rides inside `tEXt "PNG"` instead — and placed right after `IHDR` so `%PDF-` stays within the header scan window (§4.3). |
|
|
364
|
+
| `pixel-data chunks` | PNG | PNG face | Every chunk of the source image between `IHDR` and `IEND`, ancillary chunks included, copied unmodified. The reference writer takes `IHDR` as the 25 bytes after the signature and `IEND` as the last 12 bytes of the source, so a source image with bytes after `IEND` is not supported. With the HTML face the chunks sit inside the wrapper so the HTML parser skips them. |
|
|
365
|
+
| `tEXt "ZIP"` | PNG | PNG face | The 12 header bytes (length, type, keyword, NUL separator) of the archive's own `tEXt` chunk — the second one when the HTML face or the PDF face put a chunk ahead of it, the only one otherwise. Its declared length covers everything from there up to but not including the trailing chunk CRC, as a PNG chunk length always does, so the PNG decoder skips the archive — and, with the HTML face, the bootstrap and the appended data — as the data of one chunk. With the HTML face, the wrapper opened at the end of `tEXt "PNG"` closes immediately after these bytes: its content is the first chunk's CRC, the pixel-data chunks and this chunk's own header, and the prologue resumes as markup directly after the close tag. |
|
|
327
366
|
| `crc · IEND` | PNG | PNG face | The `tEXt "ZIP"` chunk's CRC, computed once the archive bytes are final (§6), followed by the empty `IEND` chunk — the last bytes of the file (PNG requires `IEND` to end the stream, which is why the PNG variants drop the end tags). |
|
|
328
367
|
|
|
329
368
|
The reader-by-reader interpretation of these regions is §4; the mechanics that keep
|
|
@@ -360,16 +399,17 @@ The binary regions are kept out of the rendered page by the wrapper tags. The
|
|
|
360
399
|
default wrapper is an HTML comment, and the HTML standard defines exactly which
|
|
361
400
|
character sequences terminate one (`-->`, and the recovery form `--!>`); the writer
|
|
362
401
|
MUST select a wrapper only after checking the bytes it must hide against that
|
|
363
|
-
wrapper's patterns (the
|
|
364
|
-
PNG payloads, §5.1), so hiding relies on
|
|
402
|
+
wrapper's patterns (the same test for every payload, with a shorter ladder for the
|
|
403
|
+
PDF and PNG payloads and one extra check for the PNG one, §5.1), so hiding relies on
|
|
404
|
+
normative parsing behavior. When no
|
|
365
405
|
wrapper fits a PDF or PNG payload, the face is dropped rather than emitted bare
|
|
366
406
|
(§5.1). Some binary content always sits *outside* a
|
|
367
407
|
wrapper: in the PNG variants, the signature, IHDR and chunk framing bytes that
|
|
368
408
|
precede the root element start tag decode to a short run of text that HTML error recovery
|
|
369
409
|
places in the (hidden) body. The backstop for all these cases is the prologue: it
|
|
370
|
-
declares `<body hidden
|
|
371
|
-
wait and error messages
|
|
372
|
-
with or without scripting.
|
|
410
|
+
declares `<body hidden>`, which only the bootstrap clears, and a stylesheet that
|
|
411
|
+
suppresses everything except the wait and error messages once the body is shown, so
|
|
412
|
+
the page comes up blank rather than showing raw bytes, with or without scripting.
|
|
373
413
|
|
|
374
414
|
The bootstrap script runs at parse time and proceeds in three stages:
|
|
375
415
|
|
|
@@ -381,8 +421,10 @@ The bootstrap script runs at parse time and proceeds in three stages:
|
|
|
381
421
|
range reading, fetching only the central directory and the entries it needs (a
|
|
382
422
|
large archive displays without downloading the ZIP region in full); otherwise it
|
|
383
423
|
downloads the whole file. When the header probe fails it falls back to page-text
|
|
384
|
-
extraction; a failure of the full download itself, past the probe,
|
|
385
|
-
|
|
424
|
+
extraction; so does a failure of the full download itself, past the probe, which
|
|
425
|
+
is why the probe leaves the document in place. A failure inside range reading,
|
|
426
|
+
past the probe, is not caught the same way: it goes to the error message. Only
|
|
427
|
+
when every applicable rung fails does the error
|
|
386
428
|
message appear, with recovery instructions that differ by variant (§2).
|
|
387
429
|
2. **Extract.** The embedded ZIP reader reads the archive through the ZIP lens
|
|
388
430
|
(§4.2) and rebuilds the page: text entries are decoded, binary entries become
|
|
@@ -432,9 +474,10 @@ A listing shows `page.pdf` first (when the PDF face is present with HTML), then
|
|
|
432
474
|
conventions below describe the reference writer rather than constraining the format —
|
|
433
475
|
readers address entries by name (§7.1):
|
|
434
476
|
|
|
435
|
-
- Entries for resources fetched from a URL carry that URL in their *comment* field
|
|
436
|
-
|
|
437
|
-
|
|
477
|
+
- Entries for resources fetched from a URL carry that URL in their *comment* field,
|
|
478
|
+
and so does each `index.html`, whose comment is the URL of the page or frame it
|
|
479
|
+
holds; a resource that came from a `data:` URL carries the literal marker `data:`
|
|
480
|
+
instead, and `manifest.json` and `page.pdf` have no comment. Comments are omitted entirely
|
|
438
481
|
from a password-protected archive, because the central directory is not encrypted
|
|
439
482
|
(§5.6).
|
|
440
483
|
- Entries whose content is already compressed (images, fonts, media, PDF) are STOREd
|
|
@@ -461,7 +504,8 @@ both. **Raw** — the EOCD declares a zero-length comment and the trailing bytes
|
|
|
461
504
|
simply outside the archive — is the default, because tools print a declared archive
|
|
462
505
|
comment on ordinary operations (§8.1), and in universal mode that comment is the
|
|
463
506
|
whole base64 recovery payload. **Declared** — the EOCD's comment length covers every
|
|
464
|
-
byte after the record
|
|
507
|
+
byte after the record, `declareAppendedData` in the reference writer — is the later
|
|
508
|
+
addition, and it is the only form some
|
|
465
509
|
readers accept at all: `java.util.zip`, and therefore Android and most JVM tooling,
|
|
466
510
|
rejects an archive with undeclared trailing bytes outright (§8.1). A writer SHOULD
|
|
467
511
|
offer both and default to raw.
|
|
@@ -495,15 +539,16 @@ compatibility appendix records real-world support (§8):
|
|
|
495
539
|
the objects themselves.
|
|
496
540
|
|
|
497
541
|
With the HTML face, the PDF document sits inside the head, wrapped in its own
|
|
498
|
-
wrapper-tag pair chosen against the
|
|
499
|
-
ladder,
|
|
500
|
-
|
|
542
|
+
wrapper-tag pair chosen against the local header and the PDF together (§5.1) — a PDF
|
|
543
|
+
containing `-->` steps the ladder, and so can the header, whose CRC-32 and size fields
|
|
544
|
+
hold arbitrary bytes — and preceded by the `page.pdf` local header so the same bytes
|
|
545
|
+
are also a ZIP entry. That dual role is why the entry MUST be STOREd and MUST NOT be
|
|
501
546
|
encrypted: a viewer reads the entry's data region directly, and any transformation
|
|
502
547
|
of it would break the face. Without the HTML face the document needs no wrapper, and
|
|
503
548
|
where it sits depends on the PNG face: alone with the ZIP face it simply starts the
|
|
504
549
|
file, at offset 0, exercising only the trailing-data tolerance; with the PNG face it is
|
|
505
550
|
the data of a `tEXt "PDF"` chunk placed right after `IHDR` (§3.1), which puts `%PDF-`
|
|
506
|
-
at
|
|
551
|
+
at offset 45 exactly, inside the window but not at its start.
|
|
507
552
|
|
|
508
553
|
### 4.4 The PNG decoder
|
|
509
554
|
|
|
@@ -525,9 +570,12 @@ chunks (ancillary by construction, their type starting lowercase):
|
|
|
525
570
|
ZIP region and the appended data. The decoder hops over all of it as the data of
|
|
526
571
|
one chunk.
|
|
527
572
|
|
|
528
|
-
|
|
529
|
-
string is Latin-1 text, and
|
|
530
|
-
`
|
|
573
|
+
The `tEXt "ZIP"` chunk carries text the PNG standard does not strictly permit: a
|
|
574
|
+
`tEXt` text string is Latin-1 text, and its payload contains NUL bytes — 78 in the
|
|
575
|
+
`png` specimen, 101 in the `png-pdf` one. The first chunk is pure printable ASCII in
|
|
576
|
+
the plain `png` specimen, where the doctype and the provenance comment are suppressed
|
|
577
|
+
and the title is escaped to character references; only the `-pdf` variants put NULs
|
|
578
|
+
in it. Decoders skip
|
|
531
579
|
ancillary chunks without inspecting their text, so this passes everywhere tested
|
|
532
580
|
(§8.1). It exercises the PNG tolerance §1.1 lists at its limit: what decoders ignore
|
|
533
581
|
in a `tEXt` chunk is text PNG does not permit.
|
|
@@ -554,7 +602,8 @@ It works in three steps:
|
|
|
554
602
|
document whose data starts with those characters. An element bearing the identifier
|
|
555
603
|
wins over a comment when both resolve; a reader that finds an id-bearing element
|
|
556
604
|
which is not one of §5.1's wrapper rungs SHOULD fall back to the comment, since the
|
|
557
|
-
`id` is then something else in the page
|
|
605
|
+
`id` is then something else in the page; the reference extractor does not, and
|
|
606
|
+
takes whatever element bears the identifier. The two placements (§3.1) need no
|
|
558
607
|
telling apart, and neither the region's position in the tree nor its depth carries
|
|
559
608
|
meaning — a document that moved the node before extraction resolves the same way,
|
|
560
609
|
which matters because the reference extractor relocates `meta` and `style` elements
|
|
@@ -593,7 +642,8 @@ It works in three steps:
|
|
|
593
642
|
256 values map to themselves — but the shortcut "any code point ≤ 255 is that
|
|
594
643
|
byte" is **not** a valid substitute: under other qualifying charsets (§2.1) code
|
|
595
644
|
points below 256 can belong to a different byte, 75 of them under `macintosh`, and
|
|
596
|
-
the shortcut would silently corrupt the region. The
|
|
645
|
+
the shortcut would silently corrupt the region. The reference extractor takes it,
|
|
646
|
+
and is correct only because it supports windows-1252 alone. The two things parsing destroyed
|
|
597
647
|
are restored from the payload: each parsed newline consumes the next 2-bit code to
|
|
598
648
|
reproduce the original byte sequence.
|
|
599
649
|
|
|
@@ -635,7 +685,10 @@ entry out on every acquisition path, including the ones that read raw bytes and
|
|
|
635
685
|
return it.
|
|
636
686
|
|
|
637
687
|
The extractor MUST verify the three checkable payload fields — byte length, newline
|
|
638
|
-
count and checksum — and fail to the error message on any mismatch.
|
|
688
|
+
count and checksum — and fail to the error message on any mismatch. It MUST also fail
|
|
689
|
+
on the unassigned newline code of step 2 rather than decode it, so that a payload
|
|
690
|
+
written against a later revision of the format is named as unsupported instead of
|
|
691
|
+
silently reconstructing the wrong bytes.
|
|
639
692
|
|
|
640
693
|
The recovered region is a complete archive but **not an offset-self-contained one**.
|
|
641
694
|
Its offsets are still absolute positions in the original file (§5.3), so every
|
|
@@ -649,7 +702,11 @@ zipfile`, adds `(attempting to process anyway)` and lists both entries, while re
|
|
|
649
702
|
that compensate silently, such as Python's `zipfile`, show no diagnostic at all. The
|
|
650
703
|
shift also puts `page.pdf` out of the
|
|
651
704
|
offset-following path: its local header lies *before* the region, so its compensated
|
|
652
|
-
offset is negative
|
|
705
|
+
offset is negative and no reader can seek to it. That offset is the header's own
|
|
706
|
+
position — which §4.3 keeps inside the first 1024 bytes — minus the region's start,
|
|
707
|
+
which lies past the whole bootstrap, so it is negative for every archive and its
|
|
708
|
+
magnitude is essentially the region's start, so it moves with the size of the inlined
|
|
709
|
+
ZIP library and no particular value should be read into it. On
|
|
653
710
|
success the shifted bytes enter the normal extraction path (§4.2).
|
|
654
711
|
|
|
655
712
|
The shift is a file offset, and a universal-mode reader has no file. It does not need
|
|
@@ -665,10 +722,13 @@ where `eocdPosition` is the EOCD record's own offset within `region`, found by s
|
|
|
665
722
|
backward for its signature the way any ZIP reader finds it. A reader arrives at the
|
|
666
723
|
same number as a ZIP library's prepended-data compensation, which derives it from the
|
|
667
724
|
record's position rather than from the buffer's end. Do not substitute
|
|
668
|
-
`region.length - 22` for `eocdPosition
|
|
669
|
-
|
|
725
|
+
`region.length - 22` for `eocdPosition`. The two are in fact equal for every recovered
|
|
726
|
+
region, which always ends at the EOCD's last byte and declares a zero-length comment
|
|
727
|
+
(§1.3), but the habit fails the moment the same code is pointed at a file rather than a
|
|
728
|
+
recovered region: there a non-empty archive comment puts bytes after the record. Zip64
|
|
729
|
+
does not — its records precede the EOCD, which stays last.
|
|
670
730
|
|
|
671
|
-
Under zip64 (§5.
|
|
731
|
+
Under zip64 (§5.7) both of those EOCD fields are the `0xFFFFFFFF` sentinel, and the
|
|
672
732
|
zip64 end of central directory record carries the real values. Take them from there,
|
|
673
733
|
using the same `eocdPosition` arithmetic against that record's own position — the
|
|
674
734
|
zip64 locator states an absolute offset in the original file, so it needs the shift
|
|
@@ -689,18 +749,49 @@ universal mode this cuts its audience in two: a charset-oblivious tool that read
|
|
|
689
749
|
raw bytes — `grep`, plain-text search — sees intact UTF-8, while any consumer that
|
|
690
750
|
honors the declared `<meta charset>` decodes it as windows-1252 and garbles
|
|
691
751
|
non-ASCII text. That includes the HTML parser itself — harmless there, because the
|
|
692
|
-
region is hidden and replaced (§4.1) — but also HTML-aware indexers.
|
|
752
|
+
region is hidden and replaced (§4.1) — but also HTML-aware indexers. In universal mode the text body
|
|
693
753
|
opens with the page title, for the same raw-byte
|
|
694
754
|
audience — it is the first *text* in the element, which is not necessarily the
|
|
695
755
|
element's first line: the reference writer's serialization puts a newline before it.
|
|
696
|
-
|
|
697
|
-
|
|
698
|
-
|
|
699
|
-
|
|
700
|
-
|
|
701
|
-
|
|
702
|
-
|
|
703
|
-
|
|
756
|
+
Outside universal mode the title is not repeated there, since the prologue's own
|
|
757
|
+
`<title>` is already readable as bytes.
|
|
758
|
+
|
|
759
|
+
The `<title>` element takes the opposite route, and so does every other piece of
|
|
760
|
+
prologue text the writer assembles itself. Character references are resolved against
|
|
761
|
+
Unicode independently of the declared encoding — in RCDATA, where the title's content
|
|
762
|
+
sits, in ordinary element text, and in attribute values alike — so the writer emits
|
|
763
|
+
every character outside printable ASCII, along with `&`, `<`, `>` and `"`, as a
|
|
764
|
+
numeric reference. Those bytes are therefore pure ASCII and the text survives the
|
|
765
|
+
single-byte declaration intact: a page titled 日本語 shows as 日本語 in the browser tab
|
|
766
|
+
and to any conforming parser. The reference writer passes the title, the canonical
|
|
767
|
+
link's `href` and the viewport value — attribute values, which is what the `"` is for —
|
|
768
|
+
through one shared escaper. Writers that emit such text raw MUST NOT do so in
|
|
769
|
+
universal mode, where the same bytes decode as mojibake.
|
|
770
|
+
|
|
771
|
+
The text body above is the deliberate exception, not an oversight: it is left as raw
|
|
772
|
+
UTF-8 because its audience reads bytes rather than parsed text. The bootstrap script
|
|
773
|
+
is the other region outside this rule, and it is outside for a harder reason — script
|
|
774
|
+
data does not resolve character references at all, so the escape must happen in the
|
|
775
|
+
JavaScript source instead (§2.1).
|
|
776
|
+
|
|
777
|
+
An implementation-defined comment (§3.1) is the third region outside the rule, and the
|
|
778
|
+
only one with no escape available at all. Comment data does not resolve character
|
|
779
|
+
references either, and unlike script data it has no second language of its own to
|
|
780
|
+
escape in: `é` written in a comment stays `é` in every reader, so the escaper
|
|
781
|
+
does not restore the character, it replaces one unreadable form with another. The
|
|
782
|
+
reference writer therefore leaves the comment's characters alone and serializes it as
|
|
783
|
+
UTF-8 with the rest of the prologue, deliberately. What it does rewrite is the
|
|
784
|
+
comment's own terminators: a space goes before the `>` of `-->` and `--!>`, before a
|
|
785
|
+
leading `>` or `->`, and after a trailing `<!-`, so the comment cannot close itself
|
|
786
|
+
(§5.1). Its audience is whoever opens the raw file in an
|
|
787
|
+
editor or runs a text tool over it, and those decode the bytes as UTF-8 whatever the
|
|
788
|
+
declaration says; only a browser's raw view, which honors the declared charset, shows
|
|
789
|
+
the text as mojibake in universal mode. The bytes are harmless to extraction, since
|
|
790
|
+
the recovery payload covers the ZIP region alone, far past the prologue. A reader MUST
|
|
791
|
+
NOT rely on decoding this comment through the declared charset, and a writer that
|
|
792
|
+
wants it readable everywhere restricts it to printable ASCII. The page's own copy of
|
|
793
|
+
the comment, inside `index.html`, is UTF-8 in a UTF-8 document and is the one the
|
|
794
|
+
displayed page and the infobar carry.
|
|
704
795
|
|
|
705
796
|
## 5. Cross-cutting mechanics
|
|
706
797
|
|
|
@@ -736,12 +827,11 @@ case-insensitively: `</XMP>` and `</Script ` close their elements just as `</xmp
|
|
|
736
827
|
as the end ones. A stored, uncompressed resource is the realistic source of an
|
|
737
828
|
upper-case one.
|
|
738
829
|
|
|
739
|
-
Every rung hides its content unconditionally
|
|
740
|
-
|
|
741
|
-
|
|
742
|
-
|
|
743
|
-
|
|
744
|
-
demoted.
|
|
830
|
+
Every rung hides its content unconditionally, which is why `<noscript>` is not one. It
|
|
831
|
+
has the right terminator, but it is the one construct whose content is raw text only
|
|
832
|
+
while scripting is enabled and markup when it is not, so on a page opened without
|
|
833
|
+
scripting the archive bytes would reach the tree builder as tags. The rungs below it
|
|
834
|
+
hide the same payloads at no extra cost, so there is nothing to weigh against that.
|
|
745
835
|
|
|
746
836
|
The order under the comment is not arbitrary. Every rung hides its content from an HTML
|
|
747
837
|
parser, but text extractors differ, and the ZIP region is large enough that the
|
|
@@ -786,7 +876,7 @@ matching, indistinguishably from the same archive on the comment rung. The same
|
|
|
786
876
|
covered a payload holding every byte value, every rung's patterns, and the near-misses
|
|
787
877
|
`]]x>`, `] ]>`, `]>` and `]]`, in both the prologue position and mid-document.
|
|
788
878
|
|
|
789
|
-
Two things about the ladder *are* required. Whatever order a writer gives the
|
|
879
|
+
Two things about the ladder *are* required. Whatever order a writer gives the eight
|
|
790
880
|
closable rungs, it MUST apply the selection test below to every rung it considers, and
|
|
791
881
|
MUST keep `<plaintext>` available as the rung of last resort: §6.2's termination
|
|
792
882
|
argument needs one rung no payload can defeat.
|
|
@@ -796,7 +886,9 @@ with (§4.5): an element rung takes it as an `id` attribute — `<script type=sf
|
|
|
796
886
|
id=sfz-data>`, `<noframes id=sfz-data>` — and the comment rung as the first characters of
|
|
797
887
|
its data, `<!--sfz-data`. The wrappers hiding the PDF and PNG faces MUST NOT carry it:
|
|
798
888
|
those payloads are found by byte structure, and a second node bearing the identifier
|
|
799
|
-
would shadow the archive.
|
|
889
|
+
would shadow the archive. For the same reason no comment ahead of the wrapper may
|
|
890
|
+
begin with those characters, the implementation-defined comment of §3.1 included,
|
|
891
|
+
since the lookup takes the first that does.
|
|
800
892
|
|
|
801
893
|
The reference writer walks the ladder from the top and takes the first rung the payload
|
|
802
894
|
does not defeat. The test it applies is the format's, and is the same for every payload;
|
|
@@ -820,16 +912,13 @@ what differs is how far the ladder goes:
|
|
|
820
912
|
data double escaped*, where `</script>` does **not** close the element. So a payload
|
|
821
913
|
can hold `<!--` and then `<script`, contain no `</script` anywhere, pass the end test —
|
|
822
914
|
and the wrapper then swallows its own end tag, the extra-data element and the rest of
|
|
823
|
-
the document. On the other
|
|
915
|
+
the document. On the other seven the start test is genuine conservatism: a nested `<!--`
|
|
824
916
|
is a parse error inside a comment but does not close it, and the raw-text rungs hold a
|
|
825
917
|
flat run of characters with no states at all, while CDATA sections do not nest. The
|
|
826
918
|
rule is uniform deliberately: the
|
|
827
919
|
exemption would save one pattern match per rung on bytes already in memory, at the
|
|
828
920
|
cost of a special case an implementer has to remember correctly about the single rung
|
|
829
|
-
where forgetting it destroys the document.
|
|
830
|
-
latitude, in the broader form "a writer MAY skip the start test outside universal
|
|
831
|
-
mode", and the reference writer's PDF and PNG faces took it and shipped the bug
|
|
832
|
-
(§8.5).
|
|
921
|
+
where forgetting it destroys the document.
|
|
833
922
|
- **The PDF and PNG payloads** apply the same two tests, for the same reason — a face
|
|
834
923
|
that took the `<script>` rung on a payload holding `<!--` and `<script` would swallow
|
|
835
924
|
the rest of the document, title, bootstrap and extra-data element included — but the
|
|
@@ -888,6 +977,19 @@ writer hides and in every variant that hides one. Every rejection restarts the b
|
|
|
888
977
|
(§6): the wrapper choice changes the bytes preceding the archive, so the archive must
|
|
889
978
|
be rewritten at its new position.
|
|
890
979
|
|
|
980
|
+
Two fields are patched after that check. The EOCD comment-length field sits at the
|
|
981
|
+
end of the ZIP region and is patched under the declared form (§6.1, step 11); the
|
|
982
|
+
writer tests the bytes around it again with the final value in place and keeps the raw
|
|
983
|
+
form when that value would complete a pattern, since the raw form is always valid. The
|
|
984
|
+
`tEXt "ZIP"` length field sits inside the pixel-data wrapper, with the fixed `tEXt`
|
|
985
|
+
type and `ZIP` keyword after it, and is written last (step 12). The header is tested
|
|
986
|
+
with the rest of the payload, the length as zeros, which cannot join a pattern; the
|
|
987
|
+
real length is big-endian, so a pattern byte in it would have to be the most
|
|
988
|
+
significant byte of the chunk's size, and the smallest byte any pattern contains, `-`
|
|
989
|
+
at 0x2D, puts that size at 0x2D000000 bytes, about 755 MB. The writer refuses to
|
|
990
|
+
build a self-extracting PNG variant whose chunk reaches that size rather than
|
|
991
|
+
re-check the field.
|
|
992
|
+
|
|
891
993
|
### 5.2 The appended-data budget
|
|
892
994
|
|
|
893
995
|
Everything the writer emits after the EOCD record MUST fit in 65535 bytes — the
|
|
@@ -902,7 +1004,8 @@ and the writer compares its total against 65535 before committing to it. The EOC
|
|
|
902
1004
|
record's own 22 bytes sit inside the window too, giving the 65557-byte figure of §1.3.
|
|
903
1005
|
|
|
904
1006
|
Only the extra-data element can outgrow the budget: it carries one 2-bit code per
|
|
905
|
-
newline sequence in the
|
|
1007
|
+
newline sequence in the recovered range — the ZIP region without its comment-length
|
|
1008
|
+
field (§4.5), CR LF counting once, for two bytes (§5.5) — so it
|
|
906
1009
|
grows with the archive. Newline bytes
|
|
907
1010
|
occur at their natural density in compressed and STOREd binary data — about two in
|
|
908
1011
|
every 256 bytes — and the codes are compressed and base64-encoded, which measures at
|
|
@@ -921,7 +1024,13 @@ element moves in front of the wrapper start tag, ahead of the archive (§3.1). R
|
|
|
921
1024
|
for it MUST be reserved before the ZIP region is written, because inserting bytes
|
|
922
1025
|
ahead of the archive would shift every offset the ZIP writer has already committed;
|
|
923
1026
|
the reservation is padded with spaces and the real payload is written into it once
|
|
924
|
-
its final size is known (§6).
|
|
1027
|
+
its final size is known (§6). Relocation is final for the build, and a relocated
|
|
1028
|
+
archive carries no appended run at all: the writer emits neither the wrapper's
|
|
1029
|
+
terminator nor the end tags, so outside the PNG face, whose tail still follows (§5.1),
|
|
1030
|
+
the file ends at the EOCD record like a plain ZIP file and the readers that reject
|
|
1031
|
+
trailing bytes open it (§8.1). The parser closes the open
|
|
1032
|
+
comment or element at end of file, and `</body></html>` are implied, so the page
|
|
1033
|
+
renders the same.
|
|
925
1034
|
|
|
926
1035
|
### 5.3 Offset bookkeeping
|
|
927
1036
|
|
|
@@ -938,9 +1047,12 @@ self-consistent:
|
|
|
938
1047
|
interpreted from the `%PDF-` header, so embedding it needs no rewriting; the writer
|
|
939
1048
|
only MUST keep the header inside the scan window (§4.3).
|
|
940
1049
|
- **PNG has no offsets, only lengths.** Each chunk declares its data length. The
|
|
941
|
-
|
|
942
|
-
can only be written once the file's final size is
|
|
943
|
-
in place at the end (§6).
|
|
1050
|
+
`tEXt "ZIP"` chunk's length covers the whole archive, and the appended data too in
|
|
1051
|
+
the variants that have it, so it can only be written once the file's final size is
|
|
1052
|
+
known, and the writer patches it in place at the end (§6). A `tEXt` chunk precedes it
|
|
1053
|
+
whenever the HTML face or the PDF face is present — carrying the prologue or the PDF
|
|
1054
|
+
document respectively — and under the PNG face alone it is the only one. Appended
|
|
1055
|
+
data follows it only under the HTML face.
|
|
944
1056
|
|
|
945
1057
|
The injected `page.pdf` central record exploits a fourth, deliberate discrepancy. It
|
|
946
1058
|
is written directly to the output stream, bypassing the ZIP writer's own byte
|
|
@@ -1068,25 +1180,28 @@ ZIP tool can verify.
|
|
|
1068
1180
|
|
|
1069
1181
|
### 5.7 zip64
|
|
1070
1182
|
|
|
1071
|
-
The archive uses the zip64 structures whenever the ordinary
|
|
1072
|
-
it: a central directory starting beyond 4 GiB (the prefix counts
|
|
1073
|
-
§5.3), a directory 4 GiB or longer, or 65535 entries or more.
|
|
1074
|
-
|
|
1075
|
-
|
|
1183
|
+
The archive uses the zip64 end of central directory structures whenever the ordinary
|
|
1184
|
+
records cannot express it: a central directory starting beyond 4 GiB (the prefix counts
|
|
1185
|
+
toward the offset, §5.3), a directory 4 GiB or longer, or 65535 entries or more. A
|
|
1186
|
+
single entry of 4 GiB or more also produces zip64 extra fields, in that entry's local
|
|
1187
|
+
and central headers, without any zip64 end of central directory record. The reference
|
|
1188
|
+
writer never requests zip64 explicitly, so it appears only when reached, and given how
|
|
1189
|
+
large that is, effectively never in a saved page.
|
|
1076
1190
|
|
|
1077
1191
|
When it is reached, the EOCD record carries the sentinel values `0xFFFF` and
|
|
1078
|
-
`0xFFFFFFFF
|
|
1079
|
-
|
|
1080
|
-
|
|
1192
|
+
`0xFFFFFFFF`, preceded by a zip64 end of central directory record and its locator.
|
|
1193
|
+
The sentinels are not selective: the writer saturates the entry counts, the directory
|
|
1194
|
+
size and the directory offset together once zip64 is emitted, whichever one of them
|
|
1195
|
+
overflowed. The `page.pdf` record injection then applies its accounting to the zip64
|
|
1196
|
+
record instead — entry counts and directory size there, and
|
|
1081
1197
|
the locator's pointer moved by the record's length — while leaving each saturated
|
|
1082
1198
|
field at its sentinel. A writer MUST NOT let the injection push a 16-bit or 32-bit
|
|
1083
1199
|
field to its sentinel value without emitting the corresponding zip64 record: a count
|
|
1084
1200
|
of `0xFFFF` sends readers looking for a zip64 record that does not exist.
|
|
1085
1201
|
|
|
1086
|
-
|
|
1087
|
-
|
|
1088
|
-
|
|
1089
|
-
the non-zip64 build.
|
|
1202
|
+
`test/sfz-harness/zip64.js` covers this: the sentinels stay, the counts and the
|
|
1203
|
+
directory size land in the zip64 record, its directory offset points at the injected
|
|
1204
|
+
record, `page.pdf` is the first record in the directory, and a reader lists every entry.
|
|
1090
1205
|
|
|
1091
1206
|
zip64 does not conflict with universal mode. Its commonest trigger, 65535 entries or
|
|
1092
1207
|
more, is reached at any archive size, and §4.5 gives the offset arithmetic for a
|
|
@@ -1094,6 +1209,40 @@ recovered region whose EOCD fields are sentinels. What universal mode cannot car
|
|
|
1094
1209
|
a ZIP region of 2^32 bytes or more, which the recovery payload's 32-bit length field
|
|
1095
1210
|
cannot express (§5.5) — a size bound, not a zip64 one.
|
|
1096
1211
|
|
|
1212
|
+
### 5.8 Entry name encoding
|
|
1213
|
+
|
|
1214
|
+
Entry names in this format are arbitrary Unicode, and how a name is decoded is an
|
|
1215
|
+
interoperability question rather than a detail.
|
|
1216
|
+
|
|
1217
|
+
The reference writer never exercises that range. Its names are a fixed prefix, an
|
|
1218
|
+
index and an extension — `index.html`, `manifest.json`, `stylesheet_0.css`,
|
|
1219
|
+
`images/1.png`, `fonts/2.woff2`, `scripts/3.js`, `frames/4/`, `page.pdf` — and the
|
|
1220
|
+
extension comes either from a table of content types or from a URL pathname, which is
|
|
1221
|
+
percent-encoded. Every name it writes is therefore ASCII, whatever the language of the
|
|
1222
|
+
captured page. That is a property of this writer, not a guarantee of the format: a
|
|
1223
|
+
conforming writer may name entries after the resources themselves, and §7.3's rule
|
|
1224
|
+
that entry names are untrusted assumes one does.
|
|
1225
|
+
|
|
1226
|
+
How a name is encoded is ZIP's own business, not this format's: bit 11 of the general
|
|
1227
|
+
purpose bit flag selects UTF-8, and its absence selects the legacy code page. This
|
|
1228
|
+
document adds two requirements to that and specifies nothing else about it.
|
|
1229
|
+
|
|
1230
|
+
**A writer MUST set bit 11 on every entry**, not only on the entries whose names need
|
|
1231
|
+
it. The two encodings agree over printable ASCII, so setting it unconditionally costs
|
|
1232
|
+
nothing, and it means no name in the archive is decoded through the legacy path at all.
|
|
1233
|
+
|
|
1234
|
+
**A reader MUST honor the flag** rather than assume one encoding, and MUST expect to
|
|
1235
|
+
meet a clear one: the hand-built `page.pdf` records (§3.1, §6) are the only ones the
|
|
1236
|
+
reference writer does not produce through its ZIP writer, and an archive may carry them
|
|
1237
|
+
with no flag set at all. That single entry is then decoded as legacy while every other
|
|
1238
|
+
name in the same file is UTF-8.
|
|
1239
|
+
Its name is ASCII, where the two encodings agree, so a correct reader sees `page.pdf`
|
|
1240
|
+
either way — but a reader that hardcodes UTF-8 on the strength of the other entries has
|
|
1241
|
+
not covered it.
|
|
1242
|
+
|
|
1243
|
+
A name is not a path. §7.3's rule that entry names are untrusted applies to the decoded
|
|
1244
|
+
name, and decoding is the step before that check, not a substitute for it.
|
|
1245
|
+
|
|
1097
1246
|
## 6. Writer algorithm
|
|
1098
1247
|
|
|
1099
1248
|
This section specifies the reference writer's build order. It is normative in the
|
|
@@ -1111,27 +1260,66 @@ buys. The cost is not evenly spread:
|
|
|
1111
1260
|
| Plus universal mode | The recovery payload, the character round trip of §5.5, the appended-data budget of §5.2, and the retry loops of §6.2. This is where the real complexity lives, and it buys opening the file from `file:` with no cooperation |
|
|
1112
1261
|
| Plus the PDF or PNG face | The header window of §4.3 or the chunk patching of §5.3, plus a second wrapper choice for the embedded payload |
|
|
1113
1262
|
|
|
1114
|
-
|
|
1115
|
-
|
|
1263
|
+
The second row already needs a rebuild when the rung changes; the third adds the rest of
|
|
1264
|
+
the retry loops and the first value computed only once the archive is final, the
|
|
1265
|
+
recovery payload; the fourth adds the second such value, the PNG chunk length and CRC
|
|
1266
|
+
(§5.4). A writer that only wants durable saved
|
|
1116
1267
|
pages can stop at the first row; the files it produces are accepted by every reader in
|
|
1117
1268
|
§8.1.
|
|
1118
1269
|
|
|
1119
1270
|
### 6.1 Build order
|
|
1120
1271
|
|
|
1121
1272
|
1. **PNG head.** With the PNG face, copy the signature and `IHDR` from the source
|
|
1122
|
-
image unchanged.
|
|
1123
|
-
`tEXt "
|
|
1124
|
-
|
|
1273
|
+
image unchanged. With the HTML face, choose the wrapper for the pixel-data payload
|
|
1274
|
+
(§5.1) and emit the `tEXt "PNG"` chunk: its 12 header bytes, the head of the
|
|
1275
|
+
prologue through `<body hidden>` as built in steps 2 and 3, the wrapper start tag,
|
|
1276
|
+
and the chunk CRC. Without the HTML face but with the PDF face, emit the
|
|
1277
|
+
`tEXt "PDF"` chunk holding the PDF document here instead, so its header falls
|
|
1278
|
+
inside the PDF scan window (§4.3). Then copy every source chunk between `IHDR` and
|
|
1279
|
+
`IEND`, write the `tEXt "ZIP"` chunk header with a zero length that step 12 patches,
|
|
1280
|
+
and, with the HTML face, the pixel-data wrapper end tag.
|
|
1125
1281
|
2. **HTML prologue.** With the HTML face, emit the doctype (omitted under the PNG
|
|
1126
|
-
face, which owns the start of the file), the root element start tag,
|
|
1127
|
-
|
|
1282
|
+
face, which owns the start of the file), the root element start tag, the
|
|
1283
|
+
`<meta charset>` required by §2.1, any comment the implementation adds — after the
|
|
1284
|
+
charset declaration, since a comment carrying the page URL has no bound and would
|
|
1285
|
+
otherwise push that declaration out of the first 1024 bytes. The doctype is the
|
|
1286
|
+
other unbounded region ahead of the declaration, copied from the saved page with its
|
|
1287
|
+
identifiers verbatim, so a writer MUST emit a minimal doctype in its place when
|
|
1288
|
+
keeping it would push the declaration past 1024 bytes.
|
|
1289
|
+
|
|
1290
|
+
**Replace it; do not truncate it, and do not drop it.** Truncation is unsafe:
|
|
1291
|
+
a cut inside a quoted identifier leaves the tokenizer in the system-identifier
|
|
1292
|
+
state, where it consumes the markup that follows until the next `>` — swallowing
|
|
1293
|
+
the root element start tag and the `data-sfz` marker on it (§1.3), so the document
|
|
1294
|
+
loses both. Dropping the doctype parses
|
|
1295
|
+
cleanly but puts the document in quirks mode, which is the mode the blank-page
|
|
1296
|
+
backstop, the wait message and the error message are then rendered under (§4.1) —
|
|
1297
|
+
the error message most of all, since it is what a reader sees precisely when
|
|
1298
|
+
nothing else has worked. A minimal doctype is 15 bytes, keeps standards mode, and
|
|
1299
|
+
costs nothing else: the extracted page is written into the document with its own
|
|
1300
|
+
doctype (§4.1), so the outer one never governs the restored page.
|
|
1301
|
+
|
|
1302
|
+
The PNG face is the exception, and nothing is available to it either way. A PNG file
|
|
1303
|
+
MUST begin with its 8-byte signature, so no doctype can precede it, and one written
|
|
1304
|
+
after the PNG head is discarded: those bytes are character data, so the parser has
|
|
1305
|
+
left its initial insertion mode and ignores a DOCTYPE token. The variant renders in
|
|
1306
|
+
quirks mode until the extracted page replaces it, whatever the writer does, so this
|
|
1307
|
+
section requires nothing about the doctype there. A writer MAY drop it, as the
|
|
1308
|
+
reference writer does, or keep it — but a kept one is content like any other, and
|
|
1309
|
+
both windows are measured from the start of the *file*, which under this face begins
|
|
1310
|
+
45 bytes before the HTML does: the signature, `IHDR`, and the chunk length, type,
|
|
1311
|
+
keyword and NUL separator. That is 45 bytes less room than the arithmetic above
|
|
1312
|
+
suggests.
|
|
1313
|
+
|
|
1314
|
+
Then the head elements (the
|
|
1128
1315
|
`<title>` and the canonical link among them), the CSS and `<body hidden>`,
|
|
1129
|
-
the wait and error messages, the optional
|
|
1316
|
+
the wait and error messages, the optional text body, and the
|
|
1130
1317
|
bootstrap script. With a password, five of those are left out: the comment, the
|
|
1131
1318
|
title, the canonical link, the text body and the entry comments of step 6 (§5.6).
|
|
1132
1319
|
With the PNG face the head of this region,
|
|
1133
1320
|
through `<body hidden>`, is the data of the `tEXt "PNG"` chunk and the remainder is
|
|
1134
|
-
emitted after the `tEXt "ZIP"` chunk header
|
|
1321
|
+
emitted after the `tEXt "ZIP"` chunk header, which step 1 has already written and
|
|
1322
|
+
step 12 only patches; with the PDF face the
|
|
1135
1323
|
region is interrupted by step 3 as well.
|
|
1136
1324
|
|
|
1137
1325
|
Whatever a writer puts in the prologue, closing every element it opens before the
|
|
@@ -1141,37 +1329,44 @@ pages can stop at the first row; the files it produces are accepted by every rea
|
|
|
1141
1329
|
3. **Embedded PDF.** With the PDF face and the HTML face, the prologue is *split*
|
|
1142
1330
|
around the PDF, which MUST come early enough for `%PDF-` to start at offset 1024
|
|
1143
1331
|
or lower (§4.3). Only what a parser needs first precedes it — the doctype, the
|
|
1144
|
-
root element
|
|
1332
|
+
root element and the charset declaration — and everything
|
|
1145
1333
|
else in the head (title, link and meta elements, the stylesheet, `<body hidden>`,
|
|
1146
|
-
the messages, the optional
|
|
1334
|
+
the messages, the optional text body) follows it. Emit the
|
|
1147
1335
|
wrapper start tag chosen for the PDF payload (§5.1), the hand-built `page.pdf`
|
|
1148
1336
|
local file header, the PDF document, the wrapper end tag, and record the local
|
|
1149
|
-
header's absolute position; then resume the prologue.
|
|
1337
|
+
header's absolute position; then resume the prologue. The reference writer's
|
|
1338
|
+
header declares version 2.0, the language encoding flag alone, method STORE, the
|
|
1339
|
+
build's modification date in DOS form, the precomputed CRC-32, the document's
|
|
1340
|
+
length as both sizes, and no extra field; its central record adds a Unix
|
|
1341
|
+
"made by" version and external attributes of a regular file, mode 0644.
|
|
1150
1342
|
|
|
1151
1343
|
The window is reachable but not structurally guaranteed, and it is the one place
|
|
1152
1344
|
where the format depends on the writer rather than on its own layout. The
|
|
1153
1345
|
irreducible part of the prefix is small: the root element start tag, the charset
|
|
1154
1346
|
declaration, the wrapper start tag and the 38-byte local file header for
|
|
1155
|
-
`page.pdf`, plus a minimal doctype —
|
|
1156
|
-
|
|
1347
|
+
`page.pdf`, plus a minimal doctype — 92 bytes in the reference layout with a
|
|
1348
|
+
`utf-8` label, 99 with `windows-1252`, and a few more with a wrapper past the first
|
|
1349
|
+
rung (§5.1). But two
|
|
1157
1350
|
regions ahead of the header have no length the format controls: the doctype, which
|
|
1158
1351
|
is copied from the saved page and carries its public and system identifiers
|
|
1159
1352
|
verbatim, and any comment the implementation chooses to write there. Real doctypes
|
|
1160
1353
|
are small; the longest in common use, XHTML 1.1 with MathML and SVG, is about 140
|
|
1161
|
-
bytes
|
|
1162
|
-
|
|
1163
|
-
|
|
1164
|
-
|
|
1165
|
-
|
|
1354
|
+
bytes, though a crafted one is bounded only by what the parser accepts. Step 2's
|
|
1355
|
+
MUST already caps the doctype, but only far enough to keep the charset declaration
|
|
1356
|
+
inside the window; this header sits further into the file, behind the wrapper tag
|
|
1357
|
+
and a 38-byte local header, so it needs the tighter bound below and the comment
|
|
1358
|
+
needs one too. A writer MUST cap them itself, keeping
|
|
1359
|
+
everything before the local file header inside the remaining budget of roughly 930
|
|
1360
|
+
bytes (about 900 with the PNG face, whose signature, `IHDR` and first chunk header
|
|
1361
|
+
take the first 45 bytes of the same window while its variant drops the 15-byte
|
|
1362
|
+
doctype in exchange), shortening, dropping or relocating that content instead of emitting a
|
|
1166
1363
|
header outside the window. A writer that places nothing of unbounded length before
|
|
1167
1364
|
the PDF block satisfies the rule by construction and needs no check at all.
|
|
1168
1365
|
|
|
1169
1366
|
The reference writer does both. Its provenance comment is emitted after the PDF
|
|
1170
1367
|
block, so the page URL it carries cannot reach the window at all, and the prefix is
|
|
1171
1368
|
measured before the header is written: when the page's own doctype would push
|
|
1172
|
-
`%PDF-` past 1024, `<!DOCTYPE html>` is emitted in its place.
|
|
1173
|
-
doctype changes the bootstrap document's rendering mode, which costs nothing here:
|
|
1174
|
-
the extracted page is written into the document with its own doctype (§4.1). A
|
|
1369
|
+
`%PDF-` past 1024, `<!DOCTYPE html>` is emitted in its place, on step 2's rule. A
|
|
1175
1370
|
writer that must keep the page doctype has to find the room elsewhere.
|
|
1176
1371
|
|
|
1177
1372
|
Without the HTML face, the PDF is simply the first thing in the file and the
|
|
@@ -1179,7 +1374,9 @@ pages can stop at the first row; the files it produces are accepted by every rea
|
|
|
1179
1374
|
4. **Reserved extra-data.** In universal mode, when a previous pass determined that
|
|
1180
1375
|
the payload must be relocated (§5.2), emit an empty `<sfz-extra-data>` element
|
|
1181
1376
|
followed by enough spaces to fill the reservation. The padding sits **outside** the
|
|
1182
|
-
element, so the element's text stays exactly the payload.
|
|
1377
|
+
element, so the element's text stays exactly the payload. With
|
|
1378
|
+
`preventAppendedData` set from the start there is still no reservation on the first
|
|
1379
|
+
pass: that pass measures the payload, and the second reserves (§6.2).
|
|
1183
1380
|
5. **Wrapper start tag** for the ZIP region, carrying the identifier (§5.1).
|
|
1184
1381
|
6. **The archive.** Create the ZIP writer, telling it the number of bytes already
|
|
1185
1382
|
written so that its offsets are absolute (§5.3). Add `index.html` first, then
|
|
@@ -1193,16 +1390,22 @@ pages can stop at the first row; the files it produces are accepted by every rea
|
|
|
1193
1390
|
8. **Close and patch.** Close the archive, then correct the end of central directory
|
|
1194
1391
|
record for the injected record: entry counts, directory size, and the zip64
|
|
1195
1392
|
record and locator when present (§5.7).
|
|
1196
|
-
9. **
|
|
1197
|
-
|
|
1198
|
-
|
|
1199
|
-
|
|
1200
|
-
|
|
1201
|
-
|
|
1393
|
+
9. **Wrapper check, then the universal payload.** With the HTML face — not only in
|
|
1394
|
+
universal mode, since any self-extracting file needs it — read back the ZIP region
|
|
1395
|
+
and check it against the current wrapper (§5.1); on a collision, restart (§6.2).
|
|
1396
|
+
Then, in universal mode only, compute the region's CRC-32 and its newline codes,
|
|
1397
|
+
build and compress the payload, and decide its placement against the budget
|
|
1398
|
+
(§5.2), restarting when an appended payload turns out not to fit. Relocation is
|
|
1399
|
+
never undone (§6.2).
|
|
1400
|
+
10. **Appended run.** Unless appended data is prevented or the payload is relocated
|
|
1401
|
+
(§5.2), emit the wrapper end tag,
|
|
1202
1402
|
the extra-data element when it is appended, and `</body></html>` — the end tags
|
|
1203
1403
|
are omitted under the PNG face, which must end with `IEND`.
|
|
1204
1404
|
11. **Fill the reservation.** In the relocated placement, write the payload into the
|
|
1205
|
-
space reserved in step 4; if it no longer fits, restart (§6.2).
|
|
1405
|
+
space reserved in step 4; if it no longer fits, restart (§6.2). Under
|
|
1406
|
+
`declareAppendedData` (§4.2), the EOCD's comment-length field is patched here too,
|
|
1407
|
+
the appended run's length now being final, unless the value would complete a
|
|
1408
|
+
pattern of the current wrapper, in which case the raw form stays (§5.1).
|
|
1206
1409
|
12. **PNG tail.** With the PNG face, patch the `tEXt "ZIP"` chunk's length field, now
|
|
1207
1410
|
that the total size is known, compute that chunk's CRC over everything from its
|
|
1208
1411
|
type to the last byte written, and append the CRC and the `IEND` chunk.
|
|
@@ -1210,8 +1413,10 @@ pages can stop at the first row; the files it produces are accepted by every rea
|
|
|
1210
1413
|
### 6.2 The retry loops
|
|
1211
1414
|
|
|
1212
1415
|
Four conditions restart the build from step 1, and each restart carries forward what
|
|
1213
|
-
the failed pass learned.
|
|
1214
|
-
|
|
1416
|
+
the failed pass learned. Nothing a restart changes reaches the entries' bytes: a writer
|
|
1417
|
+
may compress them once and copy them into every pass, rewriting only the central
|
|
1418
|
+
directory's offsets, which is what the reference writer does. The first three
|
|
1419
|
+
terminate because each of them advances a monotone quantity:
|
|
1215
1420
|
|
|
1216
1421
|
- **Wrapper collision** (§5.1): the next pass starts at the next rung of the ladder.
|
|
1217
1422
|
The ladder is finite and its last rung, `<plaintext>`, is exempt from both selection
|
|
@@ -1220,29 +1425,43 @@ monotone quantity:
|
|
|
1220
1425
|
ahead of the archive, sized at the measured payload length plus a margin.
|
|
1221
1426
|
- **Reservation too small**: relocating the payload changes the file's layout, hence
|
|
1222
1427
|
its offsets, hence the payload, which can grow past the room reserved for it. The
|
|
1223
|
-
next pass reserves the new length plus the same margin.
|
|
1224
|
-
|
|
1225
|
-
|
|
1226
|
-
|
|
1227
|
-
|
|
1228
|
-
|
|
1229
|
-
|
|
1230
|
-
|
|
1231
|
-
|
|
1232
|
-
|
|
1233
|
-
|
|
1234
|
-
|
|
1235
|
-
|
|
1236
|
-
|
|
1237
|
-
|
|
1238
|
-
|
|
1428
|
+
next pass reserves the new length plus the same margin. This restart fires only
|
|
1429
|
+
when the payload outgrew its reservation and it reserves at least that payload, so
|
|
1430
|
+
every reservation is larger than the one before and the loop cannot revisit a size.
|
|
1431
|
+
What keeps it short is the margin. Shifting the offsets changes a few of the
|
|
1432
|
+
central directory's bytes, which changes the line-ending codes, the deflate output
|
|
1433
|
+
and the base64 rounding, so the payload moves by a few quanta of 4 characters
|
|
1434
|
+
between two layouts: measured between -16 and +20 characters over archives of 8 to
|
|
1435
|
+
2000 entries. A margin smaller than that shift buys a third pass in about one build
|
|
1436
|
+
out of four. The writer reserves the measured length plus 1 % plus 32 characters,
|
|
1437
|
+
which absorbed every shift measured.
|
|
1438
|
+
|
|
1439
|
+
There is no converse of the second: a pass that reserved room never discards it,
|
|
1440
|
+
even when the relocated payload would have fit the appended window. Relocation moves
|
|
1441
|
+
the archive, which changes the offsets, which changes the payload that made the
|
|
1442
|
+
relocation necessary, so a payload lying on the 65535-byte boundary can be too large
|
|
1443
|
+
appended and small enough relocated, and a writer that dropped the reservation could
|
|
1444
|
+
rebuild the two placements forever. Relocation is therefore final (§5.2), and the file
|
|
1445
|
+
keeps at most the reservation's own margin of dead padding.
|
|
1446
|
+
|
|
1447
|
+
The fourth stands apart from the other three, and terminates trivially because it can
|
|
1448
|
+
fire only once: if the end of central directory record cannot be patched to account for
|
|
1449
|
+
the injected `page.pdf` record — its signature not where the accounting expects it —
|
|
1450
|
+
the writer rebuilds without that record rather than leave a central directory the EOCD
|
|
1451
|
+
does not count. That is the restart enforcing §5.7's requirement that the injection
|
|
1452
|
+
never leave the two disagreeing, and the rebuilt archive simply has no `page.pdf`
|
|
1453
|
+
entry.
|
|
1239
1454
|
|
|
1240
1455
|
Given identical inputs, modification date and archive time, the process is
|
|
1241
|
-
deterministic: the same page produces the same bytes, retries included.
|
|
1242
|
-
|
|
1243
|
-
|
|
1244
|
-
|
|
1245
|
-
|
|
1456
|
+
deterministic: the same page produces the same bytes, retries included. `manifest.json`
|
|
1457
|
+
records when the archive was made (§7.1), so two builds of one page at two moments
|
|
1458
|
+
differ in that entry and in the entry sizes around it. A writer that retries MUST pin
|
|
1459
|
+
the archive time across the passes of one build rather than read the clock again on
|
|
1460
|
+
each. The reference writer reads the clock once, inside the callback that emits the
|
|
1461
|
+
entries, which runs once per build; every retry reuses the entries that callback
|
|
1462
|
+
produced. Two builds of one page still read the clock twice, so its own determinism
|
|
1463
|
+
test freezes it. A consumer MUST NOT
|
|
1464
|
+
treat the byte identity of two archives of the same page as meaningful.
|
|
1246
1465
|
|
|
1247
1466
|
## 7. Consuming SingleFile archives safely
|
|
1248
1467
|
|
|
@@ -1254,8 +1473,10 @@ handles every variant of §2 without knowing which one it has.
|
|
|
1254
1473
|
|
|
1255
1474
|
- **Read through the central directory.** Locate the End Of Central Directory record
|
|
1256
1475
|
by scanning backward from the end of the file, then follow its offset. A reader that
|
|
1257
|
-
streams local headers from offset 0 will not find an archive
|
|
1258
|
-
the HTML, PDF or PNG face
|
|
1476
|
+
streams local headers from offset 0 will not find an archive in any variant that has
|
|
1477
|
+
a face, since the file then starts with the HTML, PDF or PNG face. The variant with
|
|
1478
|
+
no face is an ordinary ZIP file and streams fine (§1.2, and §8.1 measures what such
|
|
1479
|
+
readers actually do).
|
|
1259
1480
|
- **Tolerate bytes before and after the archive.** They are the other faces, not
|
|
1260
1481
|
corruption. Offsets are absolute, so no compensation is needed (§5.3).
|
|
1261
1482
|
- **Accept both forms of appended data.** The bytes after the EOCD record may be raw
|
|
@@ -1273,15 +1494,18 @@ handles every variant of §2 without knowing which one it has.
|
|
|
1273
1494
|
- **Resolve the page entry in this order.** The archive's internal layout is
|
|
1274
1495
|
implementation-defined, and two properties of the reference layout matter to a
|
|
1275
1496
|
reader. Every entry MAY sit under a single root directory, which the reference
|
|
1276
|
-
writer names
|
|
1277
|
-
`<root>/index.html`. And a page's nested
|
|
1278
|
-
their own under `frames/<n>/`, recursively,
|
|
1279
|
-
`index.html`
|
|
1497
|
+
writer names `<milliseconds since the epoch>_<tab id>/` when asked to create one
|
|
1498
|
+
(`createRootDirectory`); the page is then `<root>/index.html`. And a page's nested
|
|
1499
|
+
frames are stored as complete pages of their own under `frames/<n>/`, recursively,
|
|
1500
|
+
each with its own `index.html` and `manifest.json`, so an archive normally holds
|
|
1501
|
+
several of both and only the outermost pair is the page. The recognition test above
|
|
1502
|
+
therefore matches every frame directory too. Since neither property
|
|
1280
1503
|
is guaranteed, a reader resolves the entry point in three steps, stopping at the
|
|
1281
1504
|
first that succeeds:
|
|
1282
1505
|
|
|
1283
|
-
1. `manifest.json`
|
|
1284
|
-
|
|
1506
|
+
1. The `indexFilename` of the `manifest.json` at the smallest directory depth,
|
|
1507
|
+
resolved against that manifest's directory, when the entry exists. This is the
|
|
1508
|
+
only authoritative answer, so a writer that departs from
|
|
1285
1509
|
the reference layout SHOULD emit the manifest even though a reader MUST NOT
|
|
1286
1510
|
require it.
|
|
1287
1511
|
2. Otherwise the `index.html` entry at the smallest directory depth.
|
|
@@ -1296,13 +1520,18 @@ handles every variant of §2 without knowing which one it has.
|
|
|
1296
1520
|
URL as `originalUrl`, the title as `title`, the save time as `archiveTime` (an ISO
|
|
1297
1521
|
8601 string), the entry name of the page as `indexFilename` and the resource-to-URL
|
|
1298
1522
|
map as `resources`. The page displays without any of it, and a reader MUST NOT require
|
|
1299
|
-
the entry or any field of it. `indexFilename` names the page relative to the
|
|
1300
|
-
directory, not as a full entry name.
|
|
1523
|
+
the entry or any field of it. `indexFilename` names the page relative to the
|
|
1524
|
+
manifest's own directory, not as a full entry name. A frame's manifest carries the
|
|
1525
|
+
same `archiveTime` as the page's. The set of fields is not closed: a reader
|
|
1301
1526
|
MUST ignore what it does not recognize.
|
|
1302
1527
|
- **Expect a `page.pdf` entry whose data lies outside the archive proper** (§4.2). It
|
|
1303
1528
|
is an ordinary STORE entry at an ordinary offset, so nothing special is needed to
|
|
1304
|
-
read it, but a reader that assumes
|
|
1305
|
-
|
|
1529
|
+
read it, but a reader that assumes the entries are contiguous will reject or
|
|
1530
|
+
mislocate it: `page.pdf`'s local header is the first in the file, and the whole
|
|
1531
|
+
bootstrap lies between its data and the next one. It is never placed under the root
|
|
1532
|
+
directory: the PDF face is one document per file, so an archive holds at most one
|
|
1533
|
+
`page.pdf`, and it sits at the top level whatever `createRootDirectory` does to the
|
|
1534
|
+
other entries.
|
|
1306
1535
|
|
|
1307
1536
|
### 7.2 Modifying
|
|
1308
1537
|
|
|
@@ -1326,8 +1555,10 @@ alongside it.
|
|
|
1326
1555
|
|
|
1327
1556
|
### 7.3 Security considerations
|
|
1328
1557
|
|
|
1329
|
-
- **Entry names are untrusted.**
|
|
1330
|
-
|
|
1558
|
+
- **Entry names are untrusted.** The reference writer's names are a fixed prefix, an
|
|
1559
|
+
index and an extension (§5.8), but nothing in the format requires that, and a writer
|
|
1560
|
+
may name entries after the resources themselves. A reader MUST sanitize them before
|
|
1561
|
+
writing to a filesystem: reject absolute paths and
|
|
1331
1562
|
`..` segments, and be aware that names may be long, may collide after case folding,
|
|
1332
1563
|
and may contain characters the local filesystem rejects.
|
|
1333
1564
|
- **Declared sizes are untrusted.** Do not pre-allocate from the declared uncompressed
|
|
@@ -1338,8 +1569,9 @@ alongside it.
|
|
|
1338
1569
|
the bootstrap in a privileged one. The format's own display path replaces the
|
|
1339
1570
|
document with the extracted page, which is not an isolation boundary by itself.
|
|
1340
1571
|
- **A password protects entry contents only** (§5.6). Entry names, sizes and dates
|
|
1341
|
-
stay readable in the central directory, and
|
|
1342
|
-
|
|
1572
|
+
stay readable in the central directory, and while the reference writer's names carry
|
|
1573
|
+
no information about the resources (§5.8), another writer's may state their
|
|
1574
|
+
filenames. The PNG and PDF faces render the page regardless. A conforming writer
|
|
1343
1575
|
withholds the five fields of §5.6, the source URLs among them, but a reader MUST NOT
|
|
1344
1576
|
read their absence as protection: nothing in the format stops a writer from emitting
|
|
1345
1577
|
any of them, so an archive of unknown provenance may state every URL in the clear.
|
|
@@ -1361,7 +1593,7 @@ only if it affects the bytes the page is built from:
|
|
|
1361
1593
|
| A recovery payload field disagrees with the reconstruction — length, newline count or checksum | **MUST** fail (§4.5). The reconstruction is wrong and nothing built from it can be trusted |
|
|
1362
1594
|
| An entry's CRC-32 or AES authentication code does not match | **SHOULD** fail for that entry, and MUST NOT present a page rebuilt from it as intact |
|
|
1363
1595
|
| `page.pdf` was reconstructed from the parsed page and its CRC-32 does not match | **MUST** discard the reconstruction (§4.5). The bytes are a guess about newlines the recovery payload does not describe, and the checksum is the only thing that tests it — unlike the row above, there is no read to have gone wrong, only an inference |
|
|
1364
|
-
| Bytes before the first local file header,
|
|
1596
|
+
| Bytes outside the archive proper — before the first local file header, after the EOCD record, or between an entry's data and the next header | **MUST** tolerate: they are the other faces (§7.1). The gap in the middle is not hypothetical: with the PDF face the bootstrap lies between `page.pdf`'s data and the ZIP region |
|
|
1365
1597
|
| The appended run exceeds the 65535-byte budget (§5.2) | Not a reader's problem: if the EOCD record was found, the archive is readable. Readers MAY warn |
|
|
1366
1598
|
| A `tEXt` chunk CRC does not match, or a chunk holds bytes PNG does not permit (§4.4) | Irrelevant to extraction; a reader of the archive MAY ignore both |
|
|
1367
1599
|
| `page.pdf` is present but its data does not begin with `%PDF-` | Not an error. The entry is data like any other |
|
|
@@ -1370,9 +1602,12 @@ only if it affects the bytes the page is built from:
|
|
|
1370
1602
|
| The recovered region (universal mode) disagrees with the same bytes read directly, in the EOCD's two comment-length bytes only | Expected, not an error. A recovered region always declares a zero-length comment (§4.5), so it differs here from any archive written in the declared form (§4.2). Compare the two only up to those bytes |
|
|
1371
1603
|
| The recovered region (universal mode) disagrees with the same bytes read directly, anywhere else | The file is not well-formed, whichever side is at fault, and a reader that has both MUST NOT silently merge them or pick per entry. Prefer the direct read — it is the writer's own output, where the recovered region is a reconstruction of it — and surface the disagreement rather than displaying either as intact |
|
|
1372
1604
|
|
|
1373
|
-
Anything the format does not constrain, a reader MUST NOT reject: entries may carry
|
|
1374
|
-
|
|
1375
|
-
|
|
1605
|
+
Anything the format does not constrain, a reader MUST NOT reject: entries may carry any
|
|
1606
|
+
extra fields, timestamps or data descriptors a ZIP writer would ordinarily emit. The few
|
|
1607
|
+
this document does constrain — the `0x9901` field of an encrypted entry (§4.2), the
|
|
1608
|
+
zip64 records (§5.7), the name-encoding flag (§5.8) — say how an entry is read, not
|
|
1609
|
+
whether it is acceptable, so this row covers them too: each is something a reader meets
|
|
1610
|
+
and reads.
|
|
1376
1611
|
|
|
1377
1612
|
## 8. Appendices
|
|
1378
1613
|
|
|
@@ -1445,15 +1680,25 @@ measured except Apple's `ditto`.
|
|
|
1445
1680
|
### 8.2 Anatomy of a small archive
|
|
1446
1681
|
|
|
1447
1682
|
Offsets in `universal.sfz.html` (123077 bytes, two entries, saved from `example.com`
|
|
1448
|
-
with the §8.3 command,
|
|
1683
|
+
with the §8.3 command, on the build named there).
|
|
1449
1684
|
The layout is the *universal* row of the byte map (§3).
|
|
1450
1685
|
|
|
1686
|
+
These numbers are one capture, not a contract. Everything from the bootstrap onward
|
|
1687
|
+
moves whenever the inlined ZIP library changes size, so treat the table as an
|
|
1688
|
+
illustration of the shape and not as values to compare a file against. What *is* fixed
|
|
1689
|
+
is the set of relations between the rows — the doctype opening the file with the root
|
|
1690
|
+
element start tag immediately after it, the charset declaration immediately after that
|
|
1691
|
+
and the comment immediately after that, the identifier's twelve bytes ahead of the
|
|
1692
|
+
region, the EOCD's directory offset being an absolute file position, and the entry
|
|
1693
|
+
order. Those are checked by `test/sfz-harness/byte-map.js`, which builds an equivalent
|
|
1694
|
+
specimen without a network.
|
|
1695
|
+
|
|
1451
1696
|
| Offset | Bytes | Region |
|
|
1452
1697
|
|---|---|---|
|
|
1453
1698
|
| 0 | `<!DOCTYPE html>` | `html-prologue` begins |
|
|
1454
|
-
|
|
|
1455
|
-
|
|
|
1456
|
-
|
|
|
1699
|
+
| 15 | `<html data-sfz>` | root element start tag; the attribute is the reference implementation's own marker (§1.3) |
|
|
1700
|
+
| 30 | `<meta charset=windows-1252>` | the charset rule, inside the first 1024 bytes (§2.1) |
|
|
1701
|
+
| 57 | `<!--` … `-->` (ends at 200) | comment written by the implementation; its content is implementation-defined, but where it may appear is not (§3.1, §4.6, §5.6). It follows the charset declaration so it cannot push it out of the prescan window |
|
|
1457
1702
|
| 200 | `<title>` … `</title>` (ends at 229) | the page title, as numeric character references (§4.6) |
|
|
1458
1703
|
| 677 | `<style>` | the stylesheet of the blank-page backstop (§4.1) |
|
|
1459
1704
|
| 855 | `<body hidden>` | start of the blank-page backstop (§4.1) |
|
|
@@ -1479,7 +1724,11 @@ Generated with the command-line client running `single-file-core` against
|
|
|
1479
1724
|
results of §8.1 were measured on the 1.5.107 build of the same specimen set, which
|
|
1480
1725
|
differs only inside the prologue and so falls in the same classes: those are grouped
|
|
1481
1726
|
by whether bytes precede the archive and follow the EOCD, which no prologue change
|
|
1482
|
-
alters.
|
|
1727
|
+
alters. The declared-form results are the exception, `declareAppendedData` being later
|
|
1728
|
+
than that build (§8.5); they were measured separately on a build that has it. The
|
|
1729
|
+
specimen names carry a `.sfz.html` suffix chosen for the harness; the conventions of
|
|
1730
|
+
§2.2 are what the clients produce, not what these files are called.
|
|
1731
|
+
`--compress-content` makes the output an archive; `extract-data-from-page`
|
|
1483
1732
|
defaults to true there, so the plain variant has to switch it off:
|
|
1484
1733
|
|
|
1485
1734
|
| Specimen | Command |
|
|
@@ -1497,20 +1746,18 @@ defaults to true there, so the plain variant has to switch it off:
|
|
|
1497
1746
|
These specimens are deliberately small, and a reader tested only against them is
|
|
1498
1747
|
undertested: they are all flat archives of two or three entries. None
|
|
1499
1748
|
exercises a root directory, `frames/<n>/` nesting, a second `index.html`, a `data:`-URL
|
|
1500
|
-
entry comment, the optional text body
|
|
1749
|
+
entry comment, the optional text body (§4.6), a UTF-8 BOM, zip64
|
|
1501
1750
|
(§5.7), a payload past the 64 KB budget, or a relocated reservation with padding left
|
|
1502
1751
|
in it. Two omissions matter more than the rest, because they are the parts of §5.1 a
|
|
1503
1752
|
writer is most likely to get wrong: no specimen defeats a rung by its **start**
|
|
1504
1753
|
pattern, and none defeats one with an **upper-case** pattern. A writer that tested only
|
|
1505
|
-
end patterns, or matched them case-sensitively, produces every specimen here unchanged
|
|
1506
|
-
— and the first of those two mistakes is one the reference writer actually shipped
|
|
1507
|
-
(§8.5).
|
|
1754
|
+
end patterns, or matched them case-sensitively, produces every specimen here unchanged.
|
|
1508
1755
|
|
|
1509
1756
|
Two specimens cannot be produced from a URL alone. The **ladder** specimen, which
|
|
1510
1757
|
forces the second rung of §5.1, needs a page referencing an image whose stored bytes
|
|
1511
1758
|
contain `-->`; the archive then wraps in `<script type=sfz-data>`. The **zip64** specimen requires
|
|
1512
1759
|
an archive past the thresholds of §5.7, so it is produced by calling the writer
|
|
1513
|
-
directly with zip64 forced on the ZIP writer.
|
|
1760
|
+
directly with zip64 forced on the ZIP writer, as `test/sfz-harness/zip64.js` does.
|
|
1514
1761
|
|
|
1515
1762
|
The measurements quoted elsewhere in this document come from the same harness: the
|
|
1516
1763
|
payload growth rate of §5.2 (86 KB → 181 bytes, 283 KB → 465, 1.07 MB → 1645, 4.2 MB
|
|
@@ -1535,7 +1782,7 @@ multi-byte ones (`utf-8`, `utf-16le`, `utf-16be`, `gbk`, `gb18030`, `big5`, `euc
|
|
|
1535
1782
|
`shift_jis`, `euc-kr`, `iso-2022-jp`) decode a lone byte sequence to U+FFFD or to fewer
|
|
1536
1783
|
than 256 characters.
|
|
1537
1784
|
The reverse table each one needs ranges from 8 entries (`iso-8859-15`) to 128
|
|
1538
|
-
(`koi8-r`, `koi8-u` and `
|
|
1785
|
+
(`koi8-r`, `koi8-u`, `ibm866` and `x-user-defined`); windows-1252 needs 27.
|
|
1539
1786
|
|
|
1540
1787
|
**That the round trip is charset-independent.** The mechanism of §5.5 — parse, then
|
|
1541
1788
|
re-encode with the reverse table, restoring newlines from the 2-bit codes and NUL from
|
|
@@ -1562,11 +1809,13 @@ predicts.
|
|
|
1562
1809
|
| August 2026 | Core 1.5.108: the ZIP region carries the identifier `sfz-data` and the extractor addresses it with that instead of deducing it from its position beside `<sfz-extra-data>` (§4.5). This fixes universal extraction on the `<style type=sfz-data>` rung, where the reference extractor's own relocation of `style` elements into the head moved the region out from under the positional rule |
|
|
1563
1810
|
| August 2026 | Core 1.5.108: the recovery payload stops two bytes short of the End Of Central Directory record, excluding its comment-length field (§1.3), which lets universal-mode archives declare their appended data as the archive comment — a writer option, for `java.util.zip` and the readers that reject undeclared trailing bytes (§4.2) |
|
|
1564
1811
|
| August 2026 | Core 1.5.108: the PDF and PNG faces test a wrapper rung's start pattern as well as its end pattern, closing the same script-data escape hole the ZIP region was already guarded against — a face payload holding `<!--` and then `<script` took the `<script type=sfz-data>` rung and swallowed the rest of the document (§5.1) |
|
|
1565
|
-
| August 2026 | Core 1.5.108: the retry loop discards a relocation reservation
|
|
1812
|
+
| August 2026 | Core 1.5.108: the retry loop never discards a relocation reservation, so a payload sitting on the appended-data boundary cannot oscillate between the two placements forever (§6.2) |
|
|
1566
1813
|
| August 2026 | Core 1.5.110: a PDF or PNG face whose payload names every rung is dropped instead of written bare (§5.1). Found by nesting an archive inside itself as both faces: the fifth level exhausts the ladder, and readers then extracted the fourth level's archive — checksums intact, no way to tell (§7.4) |
|
|
1567
1814
|
| August 2026 | Core 1.5.110: a PNG face leaving the comment rung on its checksum resumes the rung search instead of taking the next rung untested (§5.1). Taking it put a payload holding `</script>` on the script rung, where its own bytes closed the wrapper 93 bytes in and left the image data, the chunk framing and the whole ZIP region to the parser |
|
|
1568
1815
|
| August 2026 | Core 1.5.110: `<svg><![CDATA[` joins the ladder above `<plaintext>` (§5.1) — the one rung whose terminator, `]]>`, real payloads rarely carry. It gives a payload naming every element rung somewhere to go that does not cost the appended-data placement, and moves the self-nesting limit from the fifth level to the sixth |
|
|
1569
1816
|
| August 2026 | Core 1.5.115: password-protected archives withhold the provenance comment and the canonical link as well (§5.6). Both wrote the page's own URL into the prologue, beside the title that was already withheld, so the address the archive was saved from stayed in the clear |
|
|
1817
|
+
| August 2026 | Core 1.5.119: the inlined ZIP library is built ASCII-only, and §2.1 now requires it of any bootstrap. Its CP437 table had been emitted as literal characters, which the page re-decoded as windows-1252, growing the table from 256 entries to 508 and shifting every lookup by 60 — so the one entry read without the UTF-8 flag, `page.pdf`, came back mangled and no archive with a PDF face extracted in any engine (§5.8) |
|
|
1818
|
+
| August 2026 | Core 1.5.120: the hand-built `page.pdf` records set the language encoding flag, like every entry the ZIP writer produces (§5.8). Its name is ASCII, so no decoded name changes; what changes is that no entry in an archive is read through CP437 any more, closing the path the 1.5.119 defect surfaced on |
|
|
1570
1819
|
|
|
1571
1820
|
This document was itself revised in August 2026, against core 1.5.108, after several
|
|
1572
1821
|
independent reviews. One of them was a reader built from this specification alone, with
|
|
@@ -1603,3 +1852,13 @@ BOM, a user override and a transport-layer charset all outranking it — narrow
|
|
|
1603
1852
|
practice, since the raw read comes first and no encoding applies to it (§2.1) — and that
|
|
1604
1853
|
the recovery payload's 32-bit length field caps the region below 2^32 bytes, with
|
|
1605
1854
|
engine string limits binding well before that (§5.5).
|
|
1855
|
+
|
|
1856
|
+
A pass in September 2026, against core 1.5.120, read the text alone first and then
|
|
1857
|
+
checked each open question against the writer. It corrected two statements about the
|
|
1858
|
+
reference writer that the code contradicted: the retry loop never discards a
|
|
1859
|
+
reservation, and the archive time is read once per build, not once per pass (§6.2).
|
|
1860
|
+
It added what only the code could say: the PNG build steps that §6.1 had skipped, the
|
|
1861
|
+
fields patched after the wrapper check and the size below which they are harmless
|
|
1862
|
+
(§5.1), the chunks the PNG face copies (§3.1), the per-frame manifests and the root
|
|
1863
|
+
directory's name (§7.1), the `page.pdf` header fields (§6.1), and the range-reading
|
|
1864
|
+
failure path (§4.1).
|