single-file-core 1.5.119 → 1.5.121

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -3,7 +3,7 @@
3
3
  **Status: draft.** This document specifies the SingleFile archive, the polyglot file
4
4
  format produced by [SingleFile](https://github.com/gildas-lormeau/SingleFile) when it
5
5
  saves a page as a ZIP archive. It is written against the reference implementation,
6
- [single-file-core](https://github.com/gildas-lormeau/single-file-core) 1.5.108
6
+ [single-file-core](https://github.com/gildas-lormeau/single-file-core) 1.5.120
7
7
  (`processors/compression/`), and every byte-level statement has been verified on
8
8
  generated specimen files.
9
9
 
@@ -62,8 +62,9 @@ little software as possible. Each way of opening the file has a simpler fallback
62
62
  is hidden by construction, so the browser displays neither the page nor raw
63
63
  archive bytes (§4.1).
64
64
  3. Renamed to `.zip`, the file opens in a ZIP tool; the page and each resource are
65
- ordinary entries. The two measured readers that refuse a self-extracting variant are
66
- named in the ranking below.
65
+ ordinary entries. Two measured readers refuse a self-extracting variant even from
66
+ seekable input; the ranking below names them. A forward-only reader refuses it too,
67
+ and is a non-goal (§1.2).
67
68
  4. Renamed to `.pdf` or `.png` (when those faces are present), the file opens in a PDF
68
69
  viewer or an image viewer.
69
70
 
@@ -103,13 +104,15 @@ Three consequences shape everything below:
103
104
 
104
105
  ### 1.2 Non-goals
105
106
 
106
- - **Forward-only ZIP parsers.** SingleFile archives require central-directory-driven reading (the
107
- entries are preceded by non-ZIP bytes). Parsers that stream local headers from
108
- offset 0 are out of scope (§7).
107
+ - **Forward-only ZIP parsers.** Every archive with a face requires central-directory-driven
108
+ reading, since the entries are then preceded by non-ZIP bytes. The variant with no other
109
+ face is an ordinary ZIP file and streams from offset 0 like any other (§8.1, class A);
110
+ parsers that require that are out of scope for the rest (§7).
109
111
  - **In-place modification by generic ZIP tools.** The face invariants are global:
110
112
  the writer picks each hiding tag only after checking the exact bytes it must hide,
111
- the recovery payload of universal mode contains a checksum of the whole ZIP
112
- region, and the PDF and PNG structures wrap the archive (§5.4). A tool
113
+ the recovery payload of universal mode contains a checksum of the ZIP region
114
+ without its comment-length field, and the PDF and PNG structures wrap the
115
+ archive (§5.4). A tool
113
116
  that adds, removes or recompresses entries invalidates them, and most rewriters drop
114
117
  the prepended and appended regions outright. A generically rewritten file keeps at
115
118
  best its ZIP face. Editing an archive means producing a new one through the writer
@@ -117,7 +120,8 @@ Three consequences shape everything below:
117
120
  - **Multi-page archives.** The reference implementation can bundle several saved
118
121
  pages into one archive behind a routing bootstrap (`multiPageArchive`). This
119
122
  version of the document specifies single-page archives only; the multi-page
120
- layout is out of scope.
123
+ layout is out of scope, and so are the regions it adds to the prologue, which this
124
+ document does not describe.
121
125
  - **Confidentiality outside the ZIP entries.** A password encrypts ZIP entry contents
122
126
  only (AES). The PDF and PNG faces render the page content and are plaintext by
123
127
  design; the writer withholds what it can without breaking a face, as described in
@@ -133,12 +137,12 @@ Three consequences shape everything below:
133
137
  | **universal mode** | The variant whose HTML face can extract the archive from the *parsed page text*, the text and comment nodes the HTML parser produced, and therefore needs no access to its own raw bytes. Named "universal" because it works from any location, including the `file:` protocol. |
134
138
  | **wrapper tag** | The HTML construct that hides a binary region from the HTML parser, `<!--`…`-->` by default (§5.1). |
135
139
  | **appended data** | Bytes after the ZIP End Of Central Directory record. Readers tolerate them within the window their EOCD scan already covers: 65557 bytes from the end of the file (the 22-byte record plus the 65535-byte maximum comment length); "the 64 KB window" refers to this. It may be left undeclared or declared as the archive comment; both forms are valid ZIP and readers MUST accept both (§4.2). The recovery payload can be computed before that choice is made because it stops two bytes short of the record, excluding its comment-length field (see *recovered range* below). |
136
- | **ZIP region** | The contiguous byte range holding the archive proper: from the first local file header the ZIP writer emitted through the last byte of the End Of Central Directory record. It spans the `zip-entries`, `pdf-central-record` (when present) and `central-directory · eocd` blocks of §3, and in the HTML variants it is exactly the content of the last wrapper. It does **not** include `pdf-local-header` or the PDF document, which sit earlier in the file. |
140
+ | **ZIP region** | The contiguous byte range holding the archive proper: from the first local file header the ZIP writer emitted through the last byte of the End Of Central Directory record. It spans the `zip-entries`, `pdf-central-record` (when present) and `central-directory · eocd` blocks of §3, and in the HTML variants it is the content of the last wrapper, exactly so on the element rungs and preceded by the `sfz-data` identifier on the comment rung, which the extractor steps over. It does **not** include `pdf-local-header` or the PDF document, which sit earlier in the file. |
137
141
  | **archive** | The *logical* ZIP file: the set of entries the central directory describes, wherever their bytes lie. This is distinct from the ZIP region above, which is a contiguous byte range. Every entry but one has its bytes inside the region; `page.pdf` is the deliberate exception, an entry of the archive whose local header and data sit before the region (§4.2). "Archive" in this document always means the logical file, "ZIP region" always the byte range, and the two differ only in the PDF-with-HTML variants. |
138
142
  | **recovered range** | What the universal extractor reproduces (§4.5): the ZIP region minus its last two bytes, the comment-length field of the End Of Central Directory record. That field is the one part of the record whose value depends on what follows the region, so leaving it out is what lets a writer decide the appended-data form after the recovery payload is final (§4.2). The extractor supplies the two bytes itself, as zeroes — the recovered range carries no comment. |
139
143
  | **reference writer** | `createArchive()` in single-file-core `processors/compression/compression.js`. |
140
144
  | **bootstrap** | The inline script in the HTML face that locates, extracts and displays the archived page. |
141
- | **`sfz` identifiers** | Three byte-level identifiers carry the `sfz` prefix: `data-sfz`, `<sfz-extra-data>` and `sfz-data`. The prefix is inherited from SingleFileZ, the browser extension the format originated in (since merged into SingleFile), and is kept unchanged for compatibility with existing files. They are wire identifiers, not the format's name. Two of the three have a role in the format: `<sfz-extra-data>` is the element carrying the recovery payload, and `sfz-data` is the identifier the universal extractor addresses the ZIP region with — an `id` attribute on the wrapper element, or the first characters of the wrapper comment's data (§4.5). `data-sfz` is a marker the reference writer happens to put on the root element; this document mentions it only where it describes bytes those files contain. |
145
+ | **`sfz` identifiers** | Three byte-level identifiers carrying the `sfz` prefix matter to this document: `data-sfz`, `<sfz-extra-data>` and `sfz-data`. A conforming file may hold others the format says nothing about — the reference bootstrap gives its own status messages `sfz`-prefixed ids, which no reader has any reason to look for. The prefix is inherited from SingleFileZ, the browser extension the format originated in (since merged into SingleFile), and is kept unchanged for compatibility with existing files. They are wire identifiers, not the format's name. Two of the three have a role in the format: `<sfz-extra-data>` is the element carrying the recovery payload, and `sfz-data` is the identifier the universal extractor addresses the ZIP region with — an `id` attribute on the wrapper element, or the first characters of the wrapper comment's data (§4.5). `data-sfz` is a marker the reference writer happens to put on the root element; this document mentions it only where it describes bytes those files contain. |
142
146
 
143
147
  ## 2. Variants: composing faces
144
148
 
@@ -164,11 +168,17 @@ include it because the clients producing those variants enable it by default, bu
164
168
  `embeddedImage` and `embeddedPdf` compose with a non-universal self-extracting file
165
169
  just as well; the extension is then `.zip.html`. The one interaction: `page.pdf` lies
166
170
  outside the ZIP region (§1.3), so the page-text extraction path does not recover it
167
- with the rest and the extractor skips that entry (§4.5); every path that reads raw
168
- bytes sees it normally. In the last three rows there is no
171
+ with the rest. The extractor therefore filters that entry out unconditionally, on
172
+ every acquisition path including the ones that read raw bytes and could return it
173
+ (§4.5). In the last three rows there is no
169
174
  HTML face, so the option does not apply. The *Specimen* column names the measured
170
175
  reference files this document cites; §8 records how to regenerate them.
171
176
 
177
+ Other writer options shape the file without adding a face: `preventAppendedData` and
178
+ `declareAppendedData` (§4.2, §5.2), `includeBOM` (§3.1), `insertTextBody` (§4.6),
179
+ `password` (§5.6), `createRootDirectory` (§7.1), and the head-element switches
180
+ `insertCanonicalLink`, `insertMetaNoIndex` and `insertMetaCSP` (§3.1).
181
+
172
182
  Notes on composition:
173
183
 
174
184
  - **Universal mode requires the HTML face** (it is a property of the bootstrap) and is
@@ -235,8 +245,8 @@ they are out of reach. Taking the three in turn:
235
245
  (`includeBOM`), where nothing depends on the declared charset.
236
246
  - A **user override** is the one no software can prevent, and the rarest.
237
247
 
238
- On `file:` URLs the bootstrap goes straight to page-text extraction, since no raw read
239
- is available there (§4.1) — but there is also no transport layer, so the first of the
248
+ On `file:` URLs the bootstrap goes straight to page-text extraction, since it attempts
249
+ no raw read there (§4.1) — but there is also no transport layer, so the first of the
240
250
  three cannot arise on the very path that depends on the charset most.
241
251
 
242
252
  The failure is safe rather than silent, which is why the precondition is worth stating
@@ -267,12 +277,39 @@ or overridden by a server still tends to be decoded the way the extractor expect
267
277
  also keeps the reverse table small, at 27 entries (§5.5).
268
278
 
269
279
  The second part is the extra-data payload, which does **not**
270
- contain the archive: it carries only what the round trip destroys, namely a checksum
271
- and the information needed to restore newline bytes, which the parser normalizes.
280
+ contain the archive: it carries only what the round trip destroys or leaves
281
+ undetermined, namely a checksum, the recovered range's length, and the information
282
+ needed to restore newline bytes, which the parser normalizes.
272
283
  The parser also replaces NUL bytes with U+FFFD; since no byte decodes to U+FFFD under
273
284
  a qualifying encoding, the extractor maps U+FFFD back to NUL unambiguously and the
274
285
  payload needs nothing for it (§5.5).
275
286
 
287
+ The declared charset governs the **whole document**, not only the regions the format
288
+ reasons about. Everything the parser reads is decoded with it, the bootstrap script
289
+ included, and the writer's own code is therefore subject to the same single-byte
290
+ decoding as the page it carries. In universal mode the bootstrap MUST contain no
291
+ character outside ASCII.
292
+
293
+ Unlike the `<title>` (§4.6), it cannot be rescued by
294
+ character references. A `<script>` element's content is script data, a tokenizer state
295
+ that does not resolve them: `&#9786;` written there stays seven literal characters and
296
+ reaches the program as seven characters. The escape has to happen one level down, in
297
+ the JavaScript source — `\u263A` rather than `☺`, an escape the language resolves when
298
+ the script is compiled, not one the HTML parser resolves when the file is read. A
299
+ minifier will undo this if allowed to, since printing the shortest form is its default
300
+ and the shortest form of `\u263A` is the literal character. A writer that assembles
301
+ the bootstrap through a minifier MUST configure it to emit ASCII only.
302
+
303
+ The consequence of getting this wrong is worse than the mojibake a raw title produces,
304
+ and that is the reason for the MUST. A garbled title is visible; a garbled string
305
+ inside the extractor is not. A lookup table is the sharpest case. Emitted as literal
306
+ characters, a CP437 table is re-decoded as windows-1252 and grows from 256 entries to
307
+ 508, shifting every lookup past the first 32 by 60 positions — and nothing about the
308
+ page looks wrong, because the damage is confined to names the table decodes. Under
309
+ §5.8 that can be a single entry, and one is enough when it is the entry the extractor
310
+ matches by name. The requirement is on the whole bootstrap rather than on any table
311
+ inside it, because a minifier does not know which strings are load-bearing.
312
+
276
313
  ### 2.2 File name conventions
277
314
 
278
315
  The reference implementation names files by variant: `.zip` (no HTML face),
@@ -281,7 +318,8 @@ conventions for humans and pickers; **readers MUST NOT rely on the file name**.
281
318
  face is discoverable from the bytes alone: PNG and PDF by their signatures, the ZIP
282
319
  face by its End Of Central Directory record, and the HTML face by an `<html` start tag
283
320
  occurring before the first local file header — inside the first `tEXt` chunk's data in
284
- the PNG variants, where the markup begins after the chunk's keyword. Inside the
321
+ the PNG variants, where the markup begins after the chunk's keyword and its NUL
322
+ separator. Inside the
285
323
  archive, the `index.html` and `manifest.json` entries mark it as a saved page; a
286
324
  reader should identify it that way (§7.1). The self-extracting variants are told apart
287
325
  the same way: only a universal file carries an `<sfz-extra-data>` element.
@@ -289,9 +327,10 @@ the same way: only a universal file carries an `<sfz-extra-data>` element.
289
327
  ## 3. The byte map
290
328
 
291
329
  Unless a row states otherwise, the layouts below are measured from specimen files
292
- saved from `example.com` (the generation commands are in §8); the
293
- oversized-payload layout is derived from the writer rules instead, because a payload
294
- over 64 KB requires an archive too large for a readable specimen. The figure below shows
330
+ saved from `example.com` (the generation commands are in §8). The relocated row covers
331
+ two cases with one layout, `preventAppendedData` and a payload over 64 KB: the first is
332
+ measured on the relocated specimen, the second derived from the writer rules, because
333
+ such a payload requires an archive too large for a readable specimen. The figure below shows
295
334
  the regions and their order; the glossary of §3.1 is the normative list, and it states
296
335
  in text everything the figure conveys.
297
336
 
@@ -309,21 +348,21 @@ face adds, then the regions the PNG face adds.
309
348
 
310
349
  | Region | Producer | Present | Contents |
311
350
  |---|---|---|---|
312
- | `html-prologue` | HTML | HTML face | Doctype, the root element start tag, an optional implementation-defined comment, `<meta charset>`, title, optional head elements (canonical link, `robots` meta, viewport, Content-Security-Policy), minimal CSS, `<body hidden>`, wait/error messages, optional table of contents, optional text body (§4.6). The leading comment, the title, the canonical link and the text body are withheld when a password is set (§5.6). In the plain variant an optional UTF-8 BOM MAY precede the doctype (`includeBOM`); universal and PNG variants never carry one. In the PNG variants the region is split: everything through `<body hidden>` is the data of the `tEXt "PNG"` chunk, while the messages, the optional table of contents and the optional text body follow the `tEXt "ZIP"` chunk header; the doctype and the leading comment are dropped. |
313
- | `bootstrap` | HTML | HTML face | One inline `<script>`: the embedded ZIP reader, the extractor, the display routine, and the content-acquisition logic (§4.1). The wrapper start tag that opens the ZIP region follows it, directly or after a relocated `extra-data`. |
314
- | `<!--` / `-->` | HTML | HTML face | The wrapper tag pair hiding a binary region from the HTML parser — comment tags by default, another pair when the hidden bytes contain `-->` (§5.1). Drawn at each opening and closing position. The close tag is absent when appended data is prevented (`preventAppendedData`, or the `<plaintext>` wrapper which cannot close): no markup follows the archive and the wrapper runs to end-of-file. That does not mean the file ends at the EOCD — the PNG face's tail still follows, inside the wrapper, where it parses as text (§5.1). |
351
+ | `html-prologue` | HTML | HTML face | Doctype, the root element start tag, `<meta charset>`, an optional implementation-defined comment, title, optional head elements (canonical link, `robots` meta, viewport, Content-Security-Policy), minimal CSS, `<body hidden>`, wait/error messages, optional text body (§4.6). The leading comment, the title, the canonical link and the text body are withheld when a password is set (§5.6). In the plain variant an optional UTF-8 BOM MAY precede the doctype (`includeBOM`); universal and PNG variants never carry one, and the reference writer ignores the option there. In the PNG variants the region is split: everything through `<body hidden>` is the data of the `tEXt "PNG"` chunk, while the messages and the optional text body follow the `tEXt "ZIP"` chunk header; the doctype and the leading comment are dropped. |
352
+ | `bootstrap` | HTML | HTML face | One inline `<script>`: the embedded ZIP reader, the extractor, the display routine, and the content-acquisition logic (§4.1). In universal mode its bytes MUST be pure ASCII, since the declared charset decodes this region like any other and character references do not apply inside script data (§2.1). The wrapper start tag that opens the ZIP region follows it, directly or after a relocated `extra-data`. |
353
+ | `<!--` / `-->` | HTML | HTML face | The wrapper tag pair hiding a binary region from the HTML parser — comment tags by default, another pair when the hidden bytes defeat them — which `-->` is only the commonest way to do, the full test being `<!--`, `--!>`, a trailing `<!-` and, for the PNG payload, a leading `>` or `->` (§5.1). Drawn at each opening and closing position. The close tag is absent whenever the recovery payload is relocated (§5.2): under `preventAppendedData`, when the payload outgrows the appended-data budget, or on the `<plaintext>` wrapper which cannot close. No markup then follows the archive and the wrapper runs to end-of-file. That does not mean the file ends at the EOCD — the PNG face's tail still follows, inside the wrapper, where it parses as text (§5.1). |
315
354
  | `zip-entries` | ZIP | always | The archive's local file headers and entry data, written by the ZIP writer. The central directory of an archive written by the reference writer lists `index.html` (the page) first, then `manifest.json` (a JSON description of the archive: original URL, title, save time, resource-to-URL map — informative; the page displays without it), then the resources; the *physical* order of the local headers inside the region is not guaranteed to match, and readers MUST NOT rely on either order — entries are addressed by name (§7.1). |
316
355
  | `central-directory · eocd` | ZIP | always | The central-directory records followed by the End Of Central Directory record. All offsets are absolute file positions (§5.3). In the HTML+PDF variants the EOCD accounts for the injected `pdf-central-record` (how the writer achieves that is §6). |
317
356
  | `extra-data` | extractor | universal | `<sfz-extra-data>` element holding the base64, deflate-compressed recovery payload (§5.5). It always sits outside the wrapper, so it parses as a real element the extractor can address. Normal placement: after the EOCD, between the wrapper close tag and the end tags. Relocated placement, used when the payload exceeds the 64 KB appended-data window or `preventAppendedData` is set: immediately before the wrapper start tag. In the relocated form the element is followed by space padding: its room is reserved before the archive is written, because the region precedes the ZIP data and resizing it would shift every central-directory offset (§6). Neither placement carries positional meaning — the extractor finds the ZIP region by identifier, not relative to this element (§4.5). |
318
- | `</body></html>` | HTML | HTML face | The end tags closing the document after the wrapper close tag. Omitted when appended data is prevented, and in the PNG variants so the file can end with the PNG tail. |
319
- | `pdf-local-header` | ZIP | PDF face with HTML | The hand-built local file header for `page.pdf` (STORE, checksum precomputed), written immediately before the PDF document so ZIP readers see an ordinary entry whose data is the PDF (§6). |
320
- | `pdf-document` | PDF | PDF face | The raw PDF bytes. With the HTML face, wrapped together with `pdf-local-header` in a wrapper tag pair inside `html-prologue`, placed so `%PDF-` starts at offset 1024 or lower — the range PDF readers search for the header, which is what lets a PDF document start after other bytes at all (§4.3). Without the HTML face the file simply *starts* with the PDF document, as prepended data the ZIP face tolerates; `page.pdf` is then not an archive entry at all — no local header, no central record. |
357
+ | `</body></html>` | HTML | HTML face | The end tags closing the document after the wrapper close tag. Omitted whenever the recovery payload is relocated (§5.2), and in the PNG variants so the file can end with the PNG tail. |
358
+ | `pdf-local-header` | ZIP | PDF face with HTML | The hand-built local file header for `page.pdf` (STORE, checksum precomputed, language encoding flag set as on every other entry — §5.8), written immediately before the PDF document so ZIP readers see an ordinary entry whose data is the PDF (§6). |
359
+ | `pdf-document` | PDF | PDF face | The raw PDF bytes. With the HTML face, wrapped together with `pdf-local-header` in a wrapper tag pair inside `html-prologue`, placed so `%PDF-` starts at offset 1024 or lower — the range PDF readers search for the header, which is what lets a PDF document start after other bytes at all (§4.3). Without the HTML face and without the PNG face, the file simply *starts* with the PDF document, as prepended data the ZIP face tolerates; `page.pdf` is then not an archive entry at all — no local header, no central record. |
321
360
  | `pdf-central-record` | ZIP | PDF face with HTML | The central-directory record for `page.pdf`, injected *before* the writer's own central directory. The start of the central directory is the one place a record can be added without moving any offset the writer already committed, and it makes `page.pdf` the first entry ZIP tools list (§6). |
322
361
  | `png-signature · IHDR` | PNG | PNG face | The 8-byte PNG signature and the `IHDR` chunk declaring the source image's dimensions — the first 33 bytes of the file. |
323
- | `tEXt "PNG"` | PNG | PNG face with HTML | The length, type and keyword bytes of the first `tEXt` chunk. Its data is `html-prologue` (with the PDF face, the embedded PDF document rides inside it too), ending with the wrapper start tag. |
324
- | `tEXt "PDF"` | PNG | PNG + PDF faces without HTML | The length, type and keyword bytes of a `tEXt` chunk whose data is the raw PDF document. Written only when the PNG and PDF faces combine without HTML — with the HTML face the PDF rides inside `tEXt "PNG"` instead — and placed right after `IHDR` so `%PDF-` stays within the header scan window (§4.3). |
325
- | `pixel-data chunks` | PNG | PNG face | The source image's image-data chunks, copied unmodified. With the HTML face they sit inside the wrapper so the HTML parser skips them. |
326
- | `tEXt "ZIP"` | PNG | PNG face | The length, type and keyword bytes of the second `tEXt` chunk. Its declared length covers everything from there up to but not including the trailing chunk CRC, as a PNG chunk length always does, so the PNG decoder skips the archive — and, with the HTML face, the bootstrap and the appended data — as the data of one chunk. With the HTML face, the wrapper opened at the end of `tEXt "PNG"` closes immediately after these bytes: its content is the first chunk's CRC, the pixel-data chunks and this chunk's own header, and the prologue resumes as markup directly after the close tag. |
362
+ | `tEXt "PNG"` | PNG | PNG face with HTML | The 12 header bytes of the first `tEXt` chunk: the 4-byte big-endian length, the type, the keyword and its NUL separator. Its data is `html-prologue` (with the PDF face, the embedded PDF document rides inside it too), ending with the wrapper start tag. |
363
+ | `tEXt "PDF"` | PNG | PNG + PDF faces without HTML | The 12 header bytes (length, type, keyword, NUL separator) of a `tEXt` chunk whose data is the raw PDF document. Written only when the PNG and PDF faces combine without HTML — with the HTML face the PDF rides inside `tEXt "PNG"` instead — and placed right after `IHDR` so `%PDF-` stays within the header scan window (§4.3). |
364
+ | `pixel-data chunks` | PNG | PNG face | Every chunk of the source image between `IHDR` and `IEND`, ancillary chunks included, copied unmodified. The reference writer takes `IHDR` as the 25 bytes after the signature and `IEND` as the last 12 bytes of the source, so a source image with bytes after `IEND` is not supported. With the HTML face the chunks sit inside the wrapper so the HTML parser skips them. |
365
+ | `tEXt "ZIP"` | PNG | PNG face | The 12 header bytes (length, type, keyword, NUL separator) of the archive's own `tEXt` chunk — the second one when the HTML face or the PDF face put a chunk ahead of it, the only one otherwise. Its declared length covers everything from there up to but not including the trailing chunk CRC, as a PNG chunk length always does, so the PNG decoder skips the archive — and, with the HTML face, the bootstrap and the appended data — as the data of one chunk. With the HTML face, the wrapper opened at the end of `tEXt "PNG"` closes immediately after these bytes: its content is the first chunk's CRC, the pixel-data chunks and this chunk's own header, and the prologue resumes as markup directly after the close tag. |
327
366
  | `crc · IEND` | PNG | PNG face | The `tEXt "ZIP"` chunk's CRC, computed once the archive bytes are final (§6), followed by the empty `IEND` chunk — the last bytes of the file (PNG requires `IEND` to end the stream, which is why the PNG variants drop the end tags). |
328
367
 
329
368
  The reader-by-reader interpretation of these regions is §4; the mechanics that keep
@@ -360,16 +399,17 @@ The binary regions are kept out of the rendered page by the wrapper tags. The
360
399
  default wrapper is an HTML comment, and the HTML standard defines exactly which
361
400
  character sequences terminate one (`-->`, and the recovery form `--!>`); the writer
362
401
  MUST select a wrapper only after checking the bytes it must hide against that
363
- wrapper's patterns (the exact rules differ between the ZIP region and the PDF and
364
- PNG payloads, §5.1), so hiding relies on normative parsing behavior. When no
402
+ wrapper's patterns (the same test for every payload, with a shorter ladder for the
403
+ PDF and PNG payloads and one extra check for the PNG one, §5.1), so hiding relies on
404
+ normative parsing behavior. When no
365
405
  wrapper fits a PDF or PNG payload, the face is dropped rather than emitted bare
366
406
  (§5.1). Some binary content always sits *outside* a
367
407
  wrapper: in the PNG variants, the signature, IHDR and chunk framing bytes that
368
408
  precede the root element start tag decode to a short run of text that HTML error recovery
369
409
  places in the (hidden) body. The backstop for all these cases is the prologue: it
370
- declares `<body hidden>` and a stylesheet that suppresses everything except the
371
- wait and error messages, so the page comes up blank rather than showing raw bytes,
372
- with or without scripting.
410
+ declares `<body hidden>`, which only the bootstrap clears, and a stylesheet that
411
+ suppresses everything except the wait and error messages once the body is shown, so
412
+ the page comes up blank rather than showing raw bytes, with or without scripting.
373
413
 
374
414
  The bootstrap script runs at parse time and proceeds in three stages:
375
415
 
@@ -381,8 +421,10 @@ The bootstrap script runs at parse time and proceeds in three stages:
381
421
  range reading, fetching only the central directory and the entries it needs (a
382
422
  large archive displays without downloading the ZIP region in full); otherwise it
383
423
  downloads the whole file. When the header probe fails it falls back to page-text
384
- extraction; a failure of the full download itself, past the probe, reveals the
385
- error message directly. Only when every applicable rung fails does the error
424
+ extraction; so does a failure of the full download itself, past the probe, which
425
+ is why the probe leaves the document in place. A failure inside range reading,
426
+ past the probe, is not caught the same way: it goes to the error message. Only
427
+ when every applicable rung fails does the error
386
428
  message appear, with recovery instructions that differ by variant (§2).
387
429
  2. **Extract.** The embedded ZIP reader reads the archive through the ZIP lens
388
430
  (§4.2) and rebuilds the page: text entries are decoded, binary entries become
@@ -432,9 +474,10 @@ A listing shows `page.pdf` first (when the PDF face is present with HTML), then
432
474
  conventions below describe the reference writer rather than constraining the format —
433
475
  readers address entries by name (§7.1):
434
476
 
435
- - Entries for resources fetched from a URL carry that URL in their *comment* field; a
436
- resource that came from a `data:` URL carries the literal marker `data:` instead,
437
- and `manifest.json` and `page.pdf` have no comment. Comments are omitted entirely
477
+ - Entries for resources fetched from a URL carry that URL in their *comment* field,
478
+ and so does each `index.html`, whose comment is the URL of the page or frame it
479
+ holds; a resource that came from a `data:` URL carries the literal marker `data:`
480
+ instead, and `manifest.json` and `page.pdf` have no comment. Comments are omitted entirely
438
481
  from a password-protected archive, because the central directory is not encrypted
439
482
  (§5.6).
440
483
  - Entries whose content is already compressed (images, fonts, media, PDF) are STOREd
@@ -461,7 +504,8 @@ both. **Raw** — the EOCD declares a zero-length comment and the trailing bytes
461
504
  simply outside the archive — is the default, because tools print a declared archive
462
505
  comment on ordinary operations (§8.1), and in universal mode that comment is the
463
506
  whole base64 recovery payload. **Declared** — the EOCD's comment length covers every
464
- byte after the record — is what older writers emitted, and it is the only form some
507
+ byte after the record, `declareAppendedData` in the reference writer — is the later
508
+ addition, and it is the only form some
465
509
  readers accept at all: `java.util.zip`, and therefore Android and most JVM tooling,
466
510
  rejects an archive with undeclared trailing bytes outright (§8.1). A writer SHOULD
467
511
  offer both and default to raw.
@@ -495,15 +539,16 @@ compatibility appendix records real-world support (§8):
495
539
  the objects themselves.
496
540
 
497
541
  With the HTML face, the PDF document sits inside the head, wrapped in its own
498
- wrapper-tag pair chosen against the PDF's bytes (a PDF containing `-->` steps the
499
- ladder, §5.1), and preceded by the `page.pdf` local header so the same bytes are also
500
- a ZIP entry. That dual role is why the entry MUST be STOREd and MUST NOT be
542
+ wrapper-tag pair chosen against the local header and the PDF together (§5.1) — a PDF
543
+ containing `-->` steps the ladder, and so can the header, whose CRC-32 and size fields
544
+ hold arbitrary bytes — and preceded by the `page.pdf` local header so the same bytes
545
+ are also a ZIP entry. That dual role is why the entry MUST be STOREd and MUST NOT be
501
546
  encrypted: a viewer reads the entry's data region directly, and any transformation
502
547
  of it would break the face. Without the HTML face the document needs no wrapper, and
503
548
  where it sits depends on the PNG face: alone with the ZIP face it simply starts the
504
549
  file, at offset 0, exercising only the trailing-data tolerance; with the PNG face it is
505
550
  the data of a `tEXt "PDF"` chunk placed right after `IHDR` (§3.1), which puts `%PDF-`
506
- at roughly offset 45, inside the window but not at its start.
551
+ at offset 45 exactly, inside the window but not at its start.
507
552
 
508
553
  ### 4.4 The PNG decoder
509
554
 
@@ -525,9 +570,12 @@ chunks (ancillary by construction, their type starting lowercase):
525
570
  ZIP region and the appended data. The decoder hops over all of it as the data of
526
571
  one chunk.
527
572
 
528
- Both chunks carry text the PNG standard does not strictly permit: a `tEXt` text
529
- string is Latin-1 text, and the payloads here contain NUL bytes — 78 in the
530
- `tEXt "ZIP"` chunk of the `png` specimen, 101 in the `png-pdf` one. Decoders skip
573
+ The `tEXt "ZIP"` chunk carries text the PNG standard does not strictly permit: a
574
+ `tEXt` text string is Latin-1 text, and its payload contains NUL bytes — 78 in the
575
+ `png` specimen, 101 in the `png-pdf` one. The first chunk is pure printable ASCII in
576
+ the plain `png` specimen, where the doctype and the provenance comment are suppressed
577
+ and the title is escaped to character references; only the `-pdf` variants put NULs
578
+ in it. Decoders skip
531
579
  ancillary chunks without inspecting their text, so this passes everywhere tested
532
580
  (§8.1). It exercises the PNG tolerance §1.1 lists at its limit: what decoders ignore
533
581
  in a `tEXt` chunk is text PNG does not permit.
@@ -554,7 +602,8 @@ It works in three steps:
554
602
  document whose data starts with those characters. An element bearing the identifier
555
603
  wins over a comment when both resolve; a reader that finds an id-bearing element
556
604
  which is not one of §5.1's wrapper rungs SHOULD fall back to the comment, since the
557
- `id` is then something else in the page. The two placements (§3.1) need no
605
+ `id` is then something else in the page; the reference extractor does not, and
606
+ takes whatever element bears the identifier. The two placements (§3.1) need no
558
607
  telling apart, and neither the region's position in the tree nor its depth carries
559
608
  meaning — a document that moved the node before extraction resolves the same way,
560
609
  which matters because the reference extractor relocates `meta` and `style` elements
@@ -593,7 +642,8 @@ It works in three steps:
593
642
  256 values map to themselves — but the shortcut "any code point ≤ 255 is that
594
643
  byte" is **not** a valid substitute: under other qualifying charsets (§2.1) code
595
644
  points below 256 can belong to a different byte, 75 of them under `macintosh`, and
596
- the shortcut would silently corrupt the region. The two things parsing destroyed
645
+ the shortcut would silently corrupt the region. The reference extractor takes it,
646
+ and is correct only because it supports windows-1252 alone. The two things parsing destroyed
597
647
  are restored from the payload: each parsed newline consumes the next 2-bit code to
598
648
  reproduce the original byte sequence.
599
649
 
@@ -635,7 +685,10 @@ entry out on every acquisition path, including the ones that read raw bytes and
635
685
  return it.
636
686
 
637
687
  The extractor MUST verify the three checkable payload fields — byte length, newline
638
- count and checksum — and fail to the error message on any mismatch.
688
+ count and checksum — and fail to the error message on any mismatch. It MUST also fail
689
+ on the unassigned newline code of step 2 rather than decode it, so that a payload
690
+ written against a later revision of the format is named as unsupported instead of
691
+ silently reconstructing the wrong bytes.
639
692
 
640
693
  The recovered region is a complete archive but **not an offset-self-contained one**.
641
694
  Its offsets are still absolute positions in the original file (§5.3), so every
@@ -649,7 +702,11 @@ zipfile`, adds `(attempting to process anyway)` and lists both entries, while re
649
702
  that compensate silently, such as Python's `zipfile`, show no diagnostic at all. The
650
703
  shift also puts `page.pdf` out of the
651
704
  offset-following path: its local header lies *before* the region, so its compensated
652
- offset is negative (−102092 in the `pdf` specimen) and no reader can seek to it. On
705
+ offset is negative and no reader can seek to it. That offset is the header's own
706
+ position — which §4.3 keeps inside the first 1024 bytes — minus the region's start,
707
+ which lies past the whole bootstrap, so it is negative for every archive and its
708
+ magnitude is essentially the region's start, so it moves with the size of the inlined
709
+ ZIP library and no particular value should be read into it. On
653
710
  success the shifted bytes enter the normal extraction path (§4.2).
654
711
 
655
712
  The shift is a file offset, and a universal-mode reader has no file. It does not need
@@ -665,10 +722,13 @@ where `eocdPosition` is the EOCD record's own offset within `region`, found by s
665
722
  backward for its signature the way any ZIP reader finds it. A reader arrives at the
666
723
  same number as a ZIP library's prepended-data compensation, which derives it from the
667
724
  record's position rather than from the buffer's end. Do not substitute
668
- `region.length - 22` for `eocdPosition`: the two are equal only when the EOCD is the
669
- last record in the region, which zip64 and a non-empty archive comment both break.
725
+ `region.length - 22` for `eocdPosition`. The two are in fact equal for every recovered
726
+ region, which always ends at the EOCD's last byte and declares a zero-length comment
727
+ (§1.3), but the habit fails the moment the same code is pointed at a file rather than a
728
+ recovered region: there a non-empty archive comment puts bytes after the record. Zip64
729
+ does not — its records precede the EOCD, which stays last.
670
730
 
671
- Under zip64 (§5.4) both of those EOCD fields are the `0xFFFFFFFF` sentinel, and the
731
+ Under zip64 (§5.7) both of those EOCD fields are the `0xFFFFFFFF` sentinel, and the
672
732
  zip64 end of central directory record carries the real values. Take them from there,
673
733
  using the same `eocdPosition` arithmetic against that record's own position — the
674
734
  zip64 locator states an absolute offset in the original file, so it needs the shift
@@ -689,18 +749,49 @@ universal mode this cuts its audience in two: a charset-oblivious tool that read
689
749
  raw bytes — `grep`, plain-text search — sees intact UTF-8, while any consumer that
690
750
  honors the declared `<meta charset>` decodes it as windows-1252 and garbles
691
751
  non-ASCII text. That includes the HTML parser itself — harmless there, because the
692
- region is hidden and replaced (§4.1) — but also HTML-aware indexers. The text body
752
+ region is hidden and replaced (§4.1) — but also HTML-aware indexers. In universal mode the text body
693
753
  opens with the page title, for the same raw-byte
694
754
  audience — it is the first *text* in the element, which is not necessarily the
695
755
  element's first line: the reference writer's serialization puts a newline before it.
696
-
697
- The `<title>` element takes the opposite route. Its content is RCDATA, where
698
- character references are resolved against Unicode independently of the declared
699
- encoding, so the writer emits every character outside printable ASCII — and `&`,
700
- `<`, `>` — as a numeric reference. The element's bytes are therefore pure ASCII and
701
- the title survives the single-byte declaration intact: a page titled 日本語 shows as
702
- 日本語 in the browser tab and to any conforming parser. Writers that emit the title
703
- raw MUST NOT do so in universal mode, where the same bytes decode as mojibake.
756
+ Outside universal mode the title is not repeated there, since the prologue's own
757
+ `<title>` is already readable as bytes.
758
+
759
+ The `<title>` element takes the opposite route, and so does every other piece of
760
+ prologue text the writer assembles itself. Character references are resolved against
761
+ Unicode independently of the declared encoding — in RCDATA, where the title's content
762
+ sits, in ordinary element text, and in attribute values alike — so the writer emits
763
+ every character outside printable ASCII, along with `&`, `<`, `>` and `"`, as a
764
+ numeric reference. Those bytes are therefore pure ASCII and the text survives the
765
+ single-byte declaration intact: a page titled 日本語 shows as 日本語 in the browser tab
766
+ and to any conforming parser. The reference writer passes the title, the canonical
767
+ link's `href` and the viewport value — attribute values, which is what the `"` is for —
768
+ through one shared escaper. Writers that emit such text raw MUST NOT do so in
769
+ universal mode, where the same bytes decode as mojibake.
770
+
771
+ The text body above is the deliberate exception, not an oversight: it is left as raw
772
+ UTF-8 because its audience reads bytes rather than parsed text. The bootstrap script
773
+ is the other region outside this rule, and it is outside for a harder reason — script
774
+ data does not resolve character references at all, so the escape must happen in the
775
+ JavaScript source instead (§2.1).
776
+
777
+ An implementation-defined comment (§3.1) is the third region outside the rule, and the
778
+ only one with no escape available at all. Comment data does not resolve character
779
+ references either, and unlike script data it has no second language of its own to
780
+ escape in: `&#233;` written in a comment stays `&#233;` in every reader, so the escaper
781
+ does not restore the character, it replaces one unreadable form with another. The
782
+ reference writer therefore leaves the comment's characters alone and serializes it as
783
+ UTF-8 with the rest of the prologue, deliberately. What it does rewrite is the
784
+ comment's own terminators: a space goes before the `>` of `-->` and `--!>`, before a
785
+ leading `>` or `->`, and after a trailing `<!-`, so the comment cannot close itself
786
+ (§5.1). Its audience is whoever opens the raw file in an
787
+ editor or runs a text tool over it, and those decode the bytes as UTF-8 whatever the
788
+ declaration says; only a browser's raw view, which honors the declared charset, shows
789
+ the text as mojibake in universal mode. The bytes are harmless to extraction, since
790
+ the recovery payload covers the ZIP region alone, far past the prologue. A reader MUST
791
+ NOT rely on decoding this comment through the declared charset, and a writer that
792
+ wants it readable everywhere restricts it to printable ASCII. The page's own copy of
793
+ the comment, inside `index.html`, is UTF-8 in a UTF-8 document and is the one the
794
+ displayed page and the infobar carry.
704
795
 
705
796
  ## 5. Cross-cutting mechanics
706
797
 
@@ -736,12 +827,11 @@ case-insensitively: `</XMP>` and `</Script ` close their elements just as `</xmp
736
827
  as the end ones. A stored, uncompressed resource is the realistic source of an
737
828
  upper-case one.
738
829
 
739
- Every rung hides its content unconditionally. `<noscript>` has the right terminator
740
- and was a rung until core 1.5.108, but it is the one construct whose content is raw
741
- text only while scripting is enabled and markup when it is not, so on a page opened
742
- without scripting the archive bytes would reach the tree builder as tags. The rungs
743
- below it hide the same payloads at no extra cost, so it was removed rather than
744
- demoted.
830
+ Every rung hides its content unconditionally, which is why `<noscript>` is not one. It
831
+ has the right terminator, but it is the one construct whose content is raw text only
832
+ while scripting is enabled and markup when it is not, so on a page opened without
833
+ scripting the archive bytes would reach the tree builder as tags. The rungs below it
834
+ hide the same payloads at no extra cost, so there is nothing to weigh against that.
745
835
 
746
836
  The order under the comment is not arbitrary. Every rung hides its content from an HTML
747
837
  parser, but text extractors differ, and the ZIP region is large enough that the
@@ -786,7 +876,7 @@ matching, indistinguishably from the same archive on the comment rung. The same
786
876
  covered a payload holding every byte value, every rung's patterns, and the near-misses
787
877
  `]]x>`, `] ]>`, `]>` and `]]`, in both the prologue position and mid-document.
788
878
 
789
- Two things about the ladder *are* required. Whatever order a writer gives the seven
879
+ Two things about the ladder *are* required. Whatever order a writer gives the eight
790
880
  closable rungs, it MUST apply the selection test below to every rung it considers, and
791
881
  MUST keep `<plaintext>` available as the rung of last resort: §6.2's termination
792
882
  argument needs one rung no payload can defeat.
@@ -796,7 +886,9 @@ with (§4.5): an element rung takes it as an `id` attribute — `<script type=sf
796
886
  id=sfz-data>`, `<noframes id=sfz-data>` — and the comment rung as the first characters of
797
887
  its data, `<!--sfz-data`. The wrappers hiding the PDF and PNG faces MUST NOT carry it:
798
888
  those payloads are found by byte structure, and a second node bearing the identifier
799
- would shadow the archive.
889
+ would shadow the archive. For the same reason no comment ahead of the wrapper may
890
+ begin with those characters, the implementation-defined comment of §3.1 included,
891
+ since the lookup takes the first that does.
800
892
 
801
893
  The reference writer walks the ladder from the top and takes the first rung the payload
802
894
  does not defeat. The test it applies is the format's, and is the same for every payload;
@@ -820,16 +912,13 @@ what differs is how far the ladder goes:
820
912
  data double escaped*, where `</script>` does **not** close the element. So a payload
821
913
  can hold `<!--` and then `<script`, contain no `</script` anywhere, pass the end test —
822
914
  and the wrapper then swallows its own end tag, the extra-data element and the rest of
823
- the document. On the other six the start test is genuine conservatism: a nested `<!--`
915
+ the document. On the other seven the start test is genuine conservatism: a nested `<!--`
824
916
  is a parse error inside a comment but does not close it, and the raw-text rungs hold a
825
917
  flat run of characters with no states at all, while CDATA sections do not nest. The
826
918
  rule is uniform deliberately: the
827
919
  exemption would save one pattern match per rung on bytes already in memory, at the
828
920
  cost of a special case an implementer has to remember correctly about the single rung
829
- where forgetting it destroys the document. An earlier draft offered exactly that
830
- latitude, in the broader form "a writer MAY skip the start test outside universal
831
- mode", and the reference writer's PDF and PNG faces took it and shipped the bug
832
- (§8.5).
921
+ where forgetting it destroys the document.
833
922
  - **The PDF and PNG payloads** apply the same two tests, for the same reason — a face
834
923
  that took the `<script>` rung on a payload holding `<!--` and `<script` would swallow
835
924
  the rest of the document, title, bootstrap and extra-data element included — but the
@@ -888,6 +977,19 @@ writer hides and in every variant that hides one. Every rejection restarts the b
888
977
  (§6): the wrapper choice changes the bytes preceding the archive, so the archive must
889
978
  be rewritten at its new position.
890
979
 
980
+ Two fields are patched after that check. The EOCD comment-length field sits at the
981
+ end of the ZIP region and is patched under the declared form (§6.1, step 11); the
982
+ writer tests the bytes around it again with the final value in place and keeps the raw
983
+ form when that value would complete a pattern, since the raw form is always valid. The
984
+ `tEXt "ZIP"` length field sits inside the pixel-data wrapper, with the fixed `tEXt`
985
+ type and `ZIP` keyword after it, and is written last (step 12). The header is tested
986
+ with the rest of the payload, the length as zeros, which cannot join a pattern; the
987
+ real length is big-endian, so a pattern byte in it would have to be the most
988
+ significant byte of the chunk's size, and the smallest byte any pattern contains, `-`
989
+ at 0x2D, puts that size at 0x2D000000 bytes, about 755 MB. The writer refuses to
990
+ build a self-extracting PNG variant whose chunk reaches that size rather than
991
+ re-check the field.
992
+
891
993
  ### 5.2 The appended-data budget
892
994
 
893
995
  Everything the writer emits after the EOCD record MUST fit in 65535 bytes — the
@@ -902,7 +1004,8 @@ and the writer compares its total against 65535 before committing to it. The EOC
902
1004
  record's own 22 bytes sit inside the window too, giving the 65557-byte figure of §1.3.
903
1005
 
904
1006
  Only the extra-data element can outgrow the budget: it carries one 2-bit code per
905
- newline sequence in the ZIP region — CR LF counts once, for two bytes (§5.5) — so it
1007
+ newline sequence in the recovered range — the ZIP region without its comment-length
1008
+ field (§4.5), CR LF counting once, for two bytes (§5.5) — so it
906
1009
  grows with the archive. Newline bytes
907
1010
  occur at their natural density in compressed and STOREd binary data — about two in
908
1011
  every 256 bytes — and the codes are compressed and base64-encoded, which measures at
@@ -921,7 +1024,13 @@ element moves in front of the wrapper start tag, ahead of the archive (§3.1). R
921
1024
  for it MUST be reserved before the ZIP region is written, because inserting bytes
922
1025
  ahead of the archive would shift every offset the ZIP writer has already committed;
923
1026
  the reservation is padded with spaces and the real payload is written into it once
924
- its final size is known (§6).
1027
+ its final size is known (§6). Relocation is final for the build, and a relocated
1028
+ archive carries no appended run at all: the writer emits neither the wrapper's
1029
+ terminator nor the end tags, so outside the PNG face, whose tail still follows (§5.1),
1030
+ the file ends at the EOCD record like a plain ZIP file and the readers that reject
1031
+ trailing bytes open it (§8.1). The parser closes the open
1032
+ comment or element at end of file, and `</body></html>` are implied, so the page
1033
+ renders the same.
925
1034
 
926
1035
  ### 5.3 Offset bookkeeping
927
1036
 
@@ -938,9 +1047,12 @@ self-consistent:
938
1047
  interpreted from the `%PDF-` header, so embedding it needs no rewriting; the writer
939
1048
  only MUST keep the header inside the scan window (§4.3).
940
1049
  - **PNG has no offsets, only lengths.** Each chunk declares its data length. The
941
- second `tEXt` chunk's length covers the whole archive and the appended data, so it
942
- can only be written once the file's final size is known, and the writer patches it
943
- in place at the end (§6).
1050
+ `tEXt "ZIP"` chunk's length covers the whole archive, and the appended data too in
1051
+ the variants that have it, so it can only be written once the file's final size is
1052
+ known, and the writer patches it in place at the end (§6). A `tEXt` chunk precedes it
1053
+ whenever the HTML face or the PDF face is present — carrying the prologue or the PDF
1054
+ document respectively — and under the PNG face alone it is the only one. Appended
1055
+ data follows it only under the HTML face.
944
1056
 
945
1057
  The injected `page.pdf` central record exploits a fourth, deliberate discrepancy. It
946
1058
  is written directly to the output stream, bypassing the ZIP writer's own byte
@@ -1068,25 +1180,28 @@ ZIP tool can verify.
1068
1180
 
1069
1181
  ### 5.7 zip64
1070
1182
 
1071
- The archive uses the zip64 structures whenever the ordinary records cannot express
1072
- it: a central directory starting beyond 4 GiB (the prefix counts toward the offset,
1073
- §5.3), a directory 4 GiB or longer, or 65535 entries or more. The reference writer
1074
- never requests zip64 explicitly, so it appears only when reached, and given how large
1075
- that is, effectively never in a saved page.
1183
+ The archive uses the zip64 end of central directory structures whenever the ordinary
1184
+ records cannot express it: a central directory starting beyond 4 GiB (the prefix counts
1185
+ toward the offset, §5.3), a directory 4 GiB or longer, or 65535 entries or more. A
1186
+ single entry of 4 GiB or more also produces zip64 extra fields, in that entry's local
1187
+ and central headers, without any zip64 end of central directory record. The reference
1188
+ writer never requests zip64 explicitly, so it appears only when reached, and given how
1189
+ large that is, effectively never in a saved page.
1076
1190
 
1077
1191
  When it is reached, the EOCD record carries the sentinel values `0xFFFF` and
1078
- `0xFFFFFFFF` in the fields that overflowed, preceded by a zip64 end of central
1079
- directory record and its locator. The `page.pdf` record injection then applies its
1080
- accounting to the zip64 record instead — entry counts and directory size there, and
1192
+ `0xFFFFFFFF`, preceded by a zip64 end of central directory record and its locator.
1193
+ The sentinels are not selective: the writer saturates the entry counts, the directory
1194
+ size and the directory offset together once zip64 is emitted, whichever one of them
1195
+ overflowed. The `page.pdf` record injection then applies its accounting to the zip64
1196
+ record instead — entry counts and directory size there, and
1081
1197
  the locator's pointer moved by the record's length — while leaving each saturated
1082
1198
  field at its sentinel. A writer MUST NOT let the injection push a 16-bit or 32-bit
1083
1199
  field to its sentinel value without emitting the corresponding zip64 record: a count
1084
1200
  of `0xFFFF` sends readers looking for a zip64 record that does not exist.
1085
1201
 
1086
- This combination has been verified on a forced-zip64 build (§8): `page.pdf` is listed
1087
- first by both Info-ZIP and the reference reader, the central directory offset in the
1088
- zip64 record points at the injected record, and extraction produces the same page as
1089
- the non-zip64 build.
1202
+ `test/sfz-harness/zip64.js` covers this: the sentinels stay, the counts and the
1203
+ directory size land in the zip64 record, its directory offset points at the injected
1204
+ record, `page.pdf` is the first record in the directory, and a reader lists every entry.
1090
1205
 
1091
1206
  zip64 does not conflict with universal mode. Its commonest trigger, 65535 entries or
1092
1207
  more, is reached at any archive size, and §4.5 gives the offset arithmetic for a
@@ -1094,6 +1209,40 @@ recovered region whose EOCD fields are sentinels. What universal mode cannot car
1094
1209
  a ZIP region of 2^32 bytes or more, which the recovery payload's 32-bit length field
1095
1210
  cannot express (§5.5) — a size bound, not a zip64 one.
1096
1211
 
1212
+ ### 5.8 Entry name encoding
1213
+
1214
+ Entry names in this format are arbitrary Unicode, and how a name is decoded is an
1215
+ interoperability question rather than a detail.
1216
+
1217
+ The reference writer never exercises that range. Its names are a fixed prefix, an
1218
+ index and an extension — `index.html`, `manifest.json`, `stylesheet_0.css`,
1219
+ `images/1.png`, `fonts/2.woff2`, `scripts/3.js`, `frames/4/`, `page.pdf` — and the
1220
+ extension comes either from a table of content types or from a URL pathname, which is
1221
+ percent-encoded. Every name it writes is therefore ASCII, whatever the language of the
1222
+ captured page. That is a property of this writer, not a guarantee of the format: a
1223
+ conforming writer may name entries after the resources themselves, and §7.3's rule
1224
+ that entry names are untrusted assumes one does.
1225
+
1226
+ How a name is encoded is ZIP's own business, not this format's: bit 11 of the general
1227
+ purpose bit flag selects UTF-8, and its absence selects the legacy code page. This
1228
+ document adds two requirements to that and specifies nothing else about it.
1229
+
1230
+ **A writer MUST set bit 11 on every entry**, not only on the entries whose names need
1231
+ it. The two encodings agree over printable ASCII, so setting it unconditionally costs
1232
+ nothing, and it means no name in the archive is decoded through the legacy path at all.
1233
+
1234
+ **A reader MUST honor the flag** rather than assume one encoding, and MUST expect to
1235
+ meet a clear one: the hand-built `page.pdf` records (§3.1, §6) are the only ones the
1236
+ reference writer does not produce through its ZIP writer, and an archive may carry them
1237
+ with no flag set at all. That single entry is then decoded as legacy while every other
1238
+ name in the same file is UTF-8.
1239
+ Its name is ASCII, where the two encodings agree, so a correct reader sees `page.pdf`
1240
+ either way — but a reader that hardcodes UTF-8 on the strength of the other entries has
1241
+ not covered it.
1242
+
1243
+ A name is not a path. §7.3's rule that entry names are untrusted applies to the decoded
1244
+ name, and decoding is the step before that check, not a substitute for it.
1245
+
1097
1246
  ## 6. Writer algorithm
1098
1247
 
1099
1248
  This section specifies the reference writer's build order. It is normative in the
@@ -1111,27 +1260,66 @@ buys. The cost is not evenly spread:
1111
1260
  | Plus universal mode | The recovery payload, the character round trip of §5.5, the appended-data budget of §5.2, and the retry loops of §6.2. This is where the real complexity lives, and it buys opening the file from `file:` with no cooperation |
1112
1261
  | Plus the PDF or PNG face | The header window of §4.3 or the chunk patching of §5.3, plus a second wrapper choice for the embedded payload |
1113
1262
 
1114
- Only the third row needs the retry loops, and only the fourth needs a value that cannot
1115
- be computed until the file is otherwise complete. A writer that only wants durable saved
1263
+ The second row already needs a rebuild when the rung changes; the third adds the rest of
1264
+ the retry loops and the first value computed only once the archive is final, the
1265
+ recovery payload; the fourth adds the second such value, the PNG chunk length and CRC
1266
+ (§5.4). A writer that only wants durable saved
1116
1267
  pages can stop at the first row; the files it produces are accepted by every reader in
1117
1268
  §8.1.
1118
1269
 
1119
1270
  ### 6.1 Build order
1120
1271
 
1121
1272
  1. **PNG head.** With the PNG face, copy the signature and `IHDR` from the source
1122
- image unchanged. Without the HTML face but with the PDF face, emit the
1123
- `tEXt "PDF"` chunk holding the PDF document here, so its header falls inside the
1124
- PDF scan window (§4.3).
1273
+ image unchanged. With the HTML face, choose the wrapper for the pixel-data payload
1274
+ (§5.1) and emit the `tEXt "PNG"` chunk: its 12 header bytes, the head of the
1275
+ prologue through `<body hidden>` as built in steps 2 and 3, the wrapper start tag,
1276
+ and the chunk CRC. Without the HTML face but with the PDF face, emit the
1277
+ `tEXt "PDF"` chunk holding the PDF document here instead, so its header falls
1278
+ inside the PDF scan window (§4.3). Then copy every source chunk between `IHDR` and
1279
+ `IEND`, write the `tEXt "ZIP"` chunk header with a zero length that step 12 patches,
1280
+ and, with the HTML face, the pixel-data wrapper end tag.
1125
1281
  2. **HTML prologue.** With the HTML face, emit the doctype (omitted under the PNG
1126
- face, which owns the start of the file), the root element start tag, any comment the
1127
- implementation adds, the `<meta charset>` required by §2.1, the head elements (the
1282
+ face, which owns the start of the file), the root element start tag, the
1283
+ `<meta charset>` required by §2.1, any comment the implementation adds — after the
1284
+ charset declaration, since a comment carrying the page URL has no bound and would
1285
+ otherwise push that declaration out of the first 1024 bytes. The doctype is the
1286
+ other unbounded region ahead of the declaration, copied from the saved page with its
1287
+ identifiers verbatim, so a writer MUST emit a minimal doctype in its place when
1288
+ keeping it would push the declaration past 1024 bytes.
1289
+
1290
+ **Replace it; do not truncate it, and do not drop it.** Truncation is unsafe:
1291
+ a cut inside a quoted identifier leaves the tokenizer in the system-identifier
1292
+ state, where it consumes the markup that follows until the next `>` — swallowing
1293
+ the root element start tag and the `data-sfz` marker on it (§1.3), so the document
1294
+ loses both. Dropping the doctype parses
1295
+ cleanly but puts the document in quirks mode, which is the mode the blank-page
1296
+ backstop, the wait message and the error message are then rendered under (§4.1) —
1297
+ the error message most of all, since it is what a reader sees precisely when
1298
+ nothing else has worked. A minimal doctype is 15 bytes, keeps standards mode, and
1299
+ costs nothing else: the extracted page is written into the document with its own
1300
+ doctype (§4.1), so the outer one never governs the restored page.
1301
+
1302
+ The PNG face is the exception, and nothing is available to it either way. A PNG file
1303
+ MUST begin with its 8-byte signature, so no doctype can precede it, and one written
1304
+ after the PNG head is discarded: those bytes are character data, so the parser has
1305
+ left its initial insertion mode and ignores a DOCTYPE token. The variant renders in
1306
+ quirks mode until the extracted page replaces it, whatever the writer does, so this
1307
+ section requires nothing about the doctype there. A writer MAY drop it, as the
1308
+ reference writer does, or keep it — but a kept one is content like any other, and
1309
+ both windows are measured from the start of the *file*, which under this face begins
1310
+ 45 bytes before the HTML does: the signature, `IHDR`, and the chunk length, type,
1311
+ keyword and NUL separator. That is 45 bytes less room than the arithmetic above
1312
+ suggests.
1313
+
1314
+ Then the head elements (the
1128
1315
  `<title>` and the canonical link among them), the CSS and `<body hidden>`,
1129
- the wait and error messages, the optional table of contents and text body, and the
1316
+ the wait and error messages, the optional text body, and the
1130
1317
  bootstrap script. With a password, five of those are left out: the comment, the
1131
1318
  title, the canonical link, the text body and the entry comments of step 6 (§5.6).
1132
1319
  With the PNG face the head of this region,
1133
1320
  through `<body hidden>`, is the data of the `tEXt "PNG"` chunk and the remainder is
1134
- emitted after the `tEXt "ZIP"` chunk header in step 12; with the PDF face the
1321
+ emitted after the `tEXt "ZIP"` chunk header, which step 1 has already written and
1322
+ step 12 only patches; with the PDF face the
1135
1323
  region is interrupted by step 3 as well.
1136
1324
 
1137
1325
  Whatever a writer puts in the prologue, closing every element it opens before the
@@ -1141,37 +1329,44 @@ pages can stop at the first row; the files it produces are accepted by every rea
1141
1329
  3. **Embedded PDF.** With the PDF face and the HTML face, the prologue is *split*
1142
1330
  around the PDF, which MUST come early enough for `%PDF-` to start at offset 1024
1143
1331
  or lower (§4.3). Only what a parser needs first precedes it — the doctype, the
1144
- root element, any leading comment and the charset declaration — and everything
1332
+ root element and the charset declaration — and everything
1145
1333
  else in the head (title, link and meta elements, the stylesheet, `<body hidden>`,
1146
- the messages, the optional table of contents and text body) follows it. Emit the
1334
+ the messages, the optional text body) follows it. Emit the
1147
1335
  wrapper start tag chosen for the PDF payload (§5.1), the hand-built `page.pdf`
1148
1336
  local file header, the PDF document, the wrapper end tag, and record the local
1149
- header's absolute position; then resume the prologue.
1337
+ header's absolute position; then resume the prologue. The reference writer's
1338
+ header declares version 2.0, the language encoding flag alone, method STORE, the
1339
+ build's modification date in DOS form, the precomputed CRC-32, the document's
1340
+ length as both sizes, and no extra field; its central record adds a Unix
1341
+ "made by" version and external attributes of a regular file, mode 0644.
1150
1342
 
1151
1343
  The window is reachable but not structurally guaranteed, and it is the one place
1152
1344
  where the format depends on the writer rather than on its own layout. The
1153
1345
  irreducible part of the prefix is small: the root element start tag, the charset
1154
1346
  declaration, the wrapper start tag and the 38-byte local file header for
1155
- `page.pdf`, plus a minimal doctype — 100 bytes in the reference layout, and a few
1156
- more with a longer charset label or a wrapper past the first rung (§5.1). But two
1347
+ `page.pdf`, plus a minimal doctype — 92 bytes in the reference layout with a
1348
+ `utf-8` label, 99 with `windows-1252`, and a few more with a wrapper past the first
1349
+ rung (§5.1). But two
1157
1350
  regions ahead of the header have no length the format controls: the doctype, which
1158
1351
  is copied from the saved page and carries its public and system identifiers
1159
1352
  verbatim, and any comment the implementation chooses to write there. Real doctypes
1160
1353
  are small; the longest in common use, XHTML 1.1 with MathML and SVG, is about 140
1161
- bytes. But nothing caps either region, so a writer MUST cap them itself, keeping
1162
- everything before the local file header inside the remaining budget of roughly 924
1163
- bytes (879 with the PNG face, whose signature, `IHDR` and first chunk header take
1164
- the first 45 bytes of the same window, and whose variant drops the doctype in
1165
- exchange), shortening, dropping or relocating that content instead of emitting a
1354
+ bytes, though a crafted one is bounded only by what the parser accepts. Step 2's
1355
+ MUST already caps the doctype, but only far enough to keep the charset declaration
1356
+ inside the window; this header sits further into the file, behind the wrapper tag
1357
+ and a 38-byte local header, so it needs the tighter bound below and the comment
1358
+ needs one too. A writer MUST cap them itself, keeping
1359
+ everything before the local file header inside the remaining budget of roughly 930
1360
+ bytes (about 900 with the PNG face, whose signature, `IHDR` and first chunk header
1361
+ take the first 45 bytes of the same window while its variant drops the 15-byte
1362
+ doctype in exchange), shortening, dropping or relocating that content instead of emitting a
1166
1363
  header outside the window. A writer that places nothing of unbounded length before
1167
1364
  the PDF block satisfies the rule by construction and needs no check at all.
1168
1365
 
1169
1366
  The reference writer does both. Its provenance comment is emitted after the PDF
1170
1367
  block, so the page URL it carries cannot reach the window at all, and the prefix is
1171
1368
  measured before the header is written: when the page's own doctype would push
1172
- `%PDF-` past 1024, `<!DOCTYPE html>` is emitted in its place. Substituting the
1173
- doctype changes the bootstrap document's rendering mode, which costs nothing here:
1174
- the extracted page is written into the document with its own doctype (§4.1). A
1369
+ `%PDF-` past 1024, `<!DOCTYPE html>` is emitted in its place, on step 2's rule. A
1175
1370
  writer that must keep the page doctype has to find the room elsewhere.
1176
1371
 
1177
1372
  Without the HTML face, the PDF is simply the first thing in the file and the
@@ -1179,7 +1374,9 @@ pages can stop at the first row; the files it produces are accepted by every rea
1179
1374
  4. **Reserved extra-data.** In universal mode, when a previous pass determined that
1180
1375
  the payload must be relocated (§5.2), emit an empty `<sfz-extra-data>` element
1181
1376
  followed by enough spaces to fill the reservation. The padding sits **outside** the
1182
- element, so the element's text stays exactly the payload.
1377
+ element, so the element's text stays exactly the payload. With
1378
+ `preventAppendedData` set from the start there is still no reservation on the first
1379
+ pass: that pass measures the payload, and the second reserves (§6.2).
1183
1380
  5. **Wrapper start tag** for the ZIP region, carrying the identifier (§5.1).
1184
1381
  6. **The archive.** Create the ZIP writer, telling it the number of bytes already
1185
1382
  written so that its offsets are absolute (§5.3). Add `index.html` first, then
@@ -1193,16 +1390,22 @@ pages can stop at the first row; the files it produces are accepted by every rea
1193
1390
  8. **Close and patch.** Close the archive, then correct the end of central directory
1194
1391
  record for the injected record: entry counts, directory size, and the zip64
1195
1392
  record and locator when present (§5.7).
1196
- 9. **Universal payload.** In universal mode, read back the ZIP region and check it
1197
- against the current wrapper (§5.1); on a collision, restart (§6.2). Otherwise
1198
- compute the region's CRC-32 and its newline codes, build and compress the payload,
1199
- and decide its placement against the budget (§5.2), restarting if the decision
1200
- differs from the current pass.
1201
- 10. **Appended run.** Unless appended data is prevented, emit the wrapper end tag,
1393
+ 9. **Wrapper check, then the universal payload.** With the HTML face — not only in
1394
+ universal mode, since any self-extracting file needs it — read back the ZIP region
1395
+ and check it against the current wrapper (§5.1); on a collision, restart (§6.2).
1396
+ Then, in universal mode only, compute the region's CRC-32 and its newline codes,
1397
+ build and compress the payload, and decide its placement against the budget
1398
+ (§5.2), restarting when an appended payload turns out not to fit. Relocation is
1399
+ never undone (§6.2).
1400
+ 10. **Appended run.** Unless appended data is prevented or the payload is relocated
1401
+ (§5.2), emit the wrapper end tag,
1202
1402
  the extra-data element when it is appended, and `</body></html>` — the end tags
1203
1403
  are omitted under the PNG face, which must end with `IEND`.
1204
1404
  11. **Fill the reservation.** In the relocated placement, write the payload into the
1205
- space reserved in step 4; if it no longer fits, restart (§6.2).
1405
+ space reserved in step 4; if it no longer fits, restart (§6.2). Under
1406
+ `declareAppendedData` (§4.2), the EOCD's comment-length field is patched here too,
1407
+ the appended run's length now being final, unless the value would complete a
1408
+ pattern of the current wrapper, in which case the raw form stays (§5.1).
1206
1409
  12. **PNG tail.** With the PNG face, patch the `tEXt "ZIP"` chunk's length field, now
1207
1410
  that the total size is known, compute that chunk's CRC over everything from its
1208
1411
  type to the last byte written, and append the CRC and the `IEND` chunk.
@@ -1210,8 +1413,10 @@ pages can stop at the first row; the files it produces are accepted by every rea
1210
1413
  ### 6.2 The retry loops
1211
1414
 
1212
1415
  Four conditions restart the build from step 1, and each restart carries forward what
1213
- the failed pass learned. The first three terminate because each of them advances a
1214
- monotone quantity:
1416
+ the failed pass learned. Nothing a restart changes reaches the entries' bytes: a writer
1417
+ may compress them once and copy them into every pass, rewriting only the central
1418
+ directory's offsets, which is what the reference writer does. The first three
1419
+ terminate because each of them advances a monotone quantity:
1215
1420
 
1216
1421
  - **Wrapper collision** (§5.1): the next pass starts at the next rung of the ladder.
1217
1422
  The ladder is finite and its last rung, `<plaintext>`, is exempt from both selection
@@ -1220,29 +1425,43 @@ monotone quantity:
1220
1425
  ahead of the archive, sized at the measured payload length plus a margin.
1221
1426
  - **Reservation too small**: relocating the payload changes the file's layout, hence
1222
1427
  its offsets, hence the payload, which can grow past the room reserved for it. The
1223
- next pass reserves the new length plus the same margin. For the loop to terminate,
1224
- each reservation MUST be strictly larger than the payload that sized it: a margin
1225
- that can round down to zero lets two passes measure the same length and reserve the
1226
- same room forever.
1227
-
1228
- The fourth is the converse of the second: a pass that reserved room but then found the
1229
- payload would fit in the appended window discards the reservation and rebuilds without
1230
- it, so the writer does not leave dead padding in the file. This step is not monotone,
1231
- and it is the only one that could keep the build alive forever: dropping the reservation
1232
- moves
1233
- the archive back, which changes the offsets, which changes the payload that made the
1234
- reservation necessary. A payload lying on the 65535-byte boundary can therefore be too
1235
- large appended and small enough relocated, and the build oscillates. A writer MUST
1236
- break that cycle: **the reservation is discarded at most once per build**, and a payload
1237
- that fits the appended window on a later pass stays in the reservation it already has.
1238
- The file then keeps at most the reservation's own margin of dead padding.
1428
+ next pass reserves the new length plus the same margin. This restart fires only
1429
+ when the payload outgrew its reservation and it reserves at least that payload, so
1430
+ every reservation is larger than the one before and the loop cannot revisit a size.
1431
+ What keeps it short is the margin. Shifting the offsets changes a few of the
1432
+ central directory's bytes, which changes the line-ending codes, the deflate output
1433
+ and the base64 rounding, so the payload moves by a few quanta of 4 characters
1434
+ between two layouts: measured between -16 and +20 characters over archives of 8 to
1435
+ 2000 entries. A margin smaller than that shift buys a third pass in about one build
1436
+ out of four. The writer reserves the measured length plus 1 % plus 32 characters,
1437
+ which absorbed every shift measured.
1438
+
1439
+ There is no converse of the second: a pass that reserved room never discards it,
1440
+ even when the relocated payload would have fit the appended window. Relocation moves
1441
+ the archive, which changes the offsets, which changes the payload that made the
1442
+ relocation necessary, so a payload lying on the 65535-byte boundary can be too large
1443
+ appended and small enough relocated, and a writer that dropped the reservation could
1444
+ rebuild the two placements forever. Relocation is therefore final (§5.2), and the file
1445
+ keeps at most the reservation's own margin of dead padding.
1446
+
1447
+ The fourth stands apart from the other three, and terminates trivially because it can
1448
+ fire only once: if the end of central directory record cannot be patched to account for
1449
+ the injected `page.pdf` record — its signature not where the accounting expects it —
1450
+ the writer rebuilds without that record rather than leave a central directory the EOCD
1451
+ does not count. That is the restart enforcing §5.7's requirement that the injection
1452
+ never leave the two disagreeing, and the rebuilt archive simply has no `page.pdf`
1453
+ entry.
1239
1454
 
1240
1455
  Given identical inputs, modification date and archive time, the process is
1241
- deterministic: the same page produces the same bytes, retries included. The archive
1242
- time is a separate input because `manifest.json` records when the archive was made
1243
- (§7.1), so two builds of one page at two moments differ in that entry and in the entry
1244
- sizes around it. A consumer MUST NOT treat the byte identity of two archives of the
1245
- same page as meaningful.
1456
+ deterministic: the same page produces the same bytes, retries included. `manifest.json`
1457
+ records when the archive was made (§7.1), so two builds of one page at two moments
1458
+ differ in that entry and in the entry sizes around it. A writer that retries MUST pin
1459
+ the archive time across the passes of one build rather than read the clock again on
1460
+ each. The reference writer reads the clock once, inside the callback that emits the
1461
+ entries, which runs once per build; every retry reuses the entries that callback
1462
+ produced. Two builds of one page still read the clock twice, so its own determinism
1463
+ test freezes it. A consumer MUST NOT
1464
+ treat the byte identity of two archives of the same page as meaningful.
1246
1465
 
1247
1466
  ## 7. Consuming SingleFile archives safely
1248
1467
 
@@ -1254,8 +1473,10 @@ handles every variant of §2 without knowing which one it has.
1254
1473
 
1255
1474
  - **Read through the central directory.** Locate the End Of Central Directory record
1256
1475
  by scanning backward from the end of the file, then follow its offset. A reader that
1257
- streams local headers from offset 0 will not find an archive: the file starts with
1258
- the HTML, PDF or PNG face (§1.2, and §8.1 measures what such readers actually do).
1476
+ streams local headers from offset 0 will not find an archive in any variant that has
1477
+ a face, since the file then starts with the HTML, PDF or PNG face. The variant with
1478
+ no face is an ordinary ZIP file and streams fine (§1.2, and §8.1 measures what such
1479
+ readers actually do).
1259
1480
  - **Tolerate bytes before and after the archive.** They are the other faces, not
1260
1481
  corruption. Offsets are absolute, so no compensation is needed (§5.3).
1261
1482
  - **Accept both forms of appended data.** The bytes after the EOCD record may be raw
@@ -1273,15 +1494,18 @@ handles every variant of §2 without knowing which one it has.
1273
1494
  - **Resolve the page entry in this order.** The archive's internal layout is
1274
1495
  implementation-defined, and two properties of the reference layout matter to a
1275
1496
  reader. Every entry MAY sit under a single root directory, which the reference
1276
- writer names from a timestamp when asked to create one; the page is then
1277
- `<root>/index.html`. And a page's nested frames are stored as complete pages of
1278
- their own under `frames/<n>/`, recursively, so an archive normally holds several
1279
- `index.html` entries and only the outermost one is the page. Since neither property
1497
+ writer names `<milliseconds since the epoch>_<tab id>/` when asked to create one
1498
+ (`createRootDirectory`); the page is then `<root>/index.html`. And a page's nested
1499
+ frames are stored as complete pages of their own under `frames/<n>/`, recursively,
1500
+ each with its own `index.html` and `manifest.json`, so an archive normally holds
1501
+ several of both and only the outermost pair is the page. The recognition test above
1502
+ therefore matches every frame directory too. Since neither property
1280
1503
  is guaranteed, a reader resolves the entry point in three steps, stopping at the
1281
1504
  first that succeeds:
1282
1505
 
1283
- 1. `manifest.json`'s `indexFilename`, resolved against the root directory, when the
1284
- entry exists. This is the only authoritative answer, so a writer that departs from
1506
+ 1. The `indexFilename` of the `manifest.json` at the smallest directory depth,
1507
+ resolved against that manifest's directory, when the entry exists. This is the
1508
+ only authoritative answer, so a writer that departs from
1285
1509
  the reference layout SHOULD emit the manifest even though a reader MUST NOT
1286
1510
  require it.
1287
1511
  2. Otherwise the `index.html` entry at the smallest directory depth.
@@ -1296,13 +1520,18 @@ handles every variant of §2 without knowing which one it has.
1296
1520
  URL as `originalUrl`, the title as `title`, the save time as `archiveTime` (an ISO
1297
1521
  8601 string), the entry name of the page as `indexFilename` and the resource-to-URL
1298
1522
  map as `resources`. The page displays without any of it, and a reader MUST NOT require
1299
- the entry or any field of it. `indexFilename` names the page relative to the root
1300
- directory, not as a full entry name. The set of fields is not closed: a reader
1523
+ the entry or any field of it. `indexFilename` names the page relative to the
1524
+ manifest's own directory, not as a full entry name. A frame's manifest carries the
1525
+ same `archiveTime` as the page's. The set of fields is not closed: a reader
1301
1526
  MUST ignore what it does not recognize.
1302
1527
  - **Expect a `page.pdf` entry whose data lies outside the archive proper** (§4.2). It
1303
1528
  is an ordinary STORE entry at an ordinary offset, so nothing special is needed to
1304
- read it, but a reader that assumes every entry sits between the first local header
1305
- and the central directory will reject or mislocate it.
1529
+ read it, but a reader that assumes the entries are contiguous will reject or
1530
+ mislocate it: `page.pdf`'s local header is the first in the file, and the whole
1531
+ bootstrap lies between its data and the next one. It is never placed under the root
1532
+ directory: the PDF face is one document per file, so an archive holds at most one
1533
+ `page.pdf`, and it sits at the top level whatever `createRootDirectory` does to the
1534
+ other entries.
1306
1535
 
1307
1536
  ### 7.2 Modifying
1308
1537
 
@@ -1326,8 +1555,10 @@ alongside it.
1326
1555
 
1327
1556
  ### 7.3 Security considerations
1328
1557
 
1329
- - **Entry names are untrusted.** They derive from a captured page's resource URLs. A
1330
- reader MUST sanitize them before writing to a filesystem: reject absolute paths and
1558
+ - **Entry names are untrusted.** The reference writer's names are a fixed prefix, an
1559
+ index and an extension (§5.8), but nothing in the format requires that, and a writer
1560
+ may name entries after the resources themselves. A reader MUST sanitize them before
1561
+ writing to a filesystem: reject absolute paths and
1331
1562
  `..` segments, and be aware that names may be long, may collide after case folding,
1332
1563
  and may contain characters the local filesystem rejects.
1333
1564
  - **Declared sizes are untrusted.** Do not pre-allocate from the declared uncompressed
@@ -1338,8 +1569,9 @@ alongside it.
1338
1569
  the bootstrap in a privileged one. The format's own display path replaces the
1339
1570
  document with the extracted page, which is not an isolation boundary by itself.
1340
1571
  - **A password protects entry contents only** (§5.6). Entry names, sizes and dates
1341
- stay readable in the central directory, and a name commonly states the resource's
1342
- filename. The PNG and PDF faces render the page regardless. A conforming writer
1572
+ stay readable in the central directory, and while the reference writer's names carry
1573
+ no information about the resources (§5.8), another writer's may state their
1574
+ filenames. The PNG and PDF faces render the page regardless. A conforming writer
1343
1575
  withholds the five fields of §5.6, the source URLs among them, but a reader MUST NOT
1344
1576
  read their absence as protection: nothing in the format stops a writer from emitting
1345
1577
  any of them, so an archive of unknown provenance may state every URL in the clear.
@@ -1361,7 +1593,7 @@ only if it affects the bytes the page is built from:
1361
1593
  | A recovery payload field disagrees with the reconstruction — length, newline count or checksum | **MUST** fail (§4.5). The reconstruction is wrong and nothing built from it can be trusted |
1362
1594
  | An entry's CRC-32 or AES authentication code does not match | **SHOULD** fail for that entry, and MUST NOT present a page rebuilt from it as intact |
1363
1595
  | `page.pdf` was reconstructed from the parsed page and its CRC-32 does not match | **MUST** discard the reconstruction (§4.5). The bytes are a guess about newlines the recovery payload does not describe, and the checksum is the only thing that tests it — unlike the row above, there is no read to have gone wrong, only an inference |
1364
- | Bytes before the first local file header, or after the EOCD record | **MUST** tolerate: they are the other faces (§7.1) |
1596
+ | Bytes outside the archive proper — before the first local file header, after the EOCD record, or between an entry's data and the next header | **MUST** tolerate: they are the other faces (§7.1). The gap in the middle is not hypothetical: with the PDF face the bootstrap lies between `page.pdf`'s data and the ZIP region |
1365
1597
  | The appended run exceeds the 65535-byte budget (§5.2) | Not a reader's problem: if the EOCD record was found, the archive is readable. Readers MAY warn |
1366
1598
  | A `tEXt` chunk CRC does not match, or a chunk holds bytes PNG does not permit (§4.4) | Irrelevant to extraction; a reader of the archive MAY ignore both |
1367
1599
  | `page.pdf` is present but its data does not begin with `%PDF-` | Not an error. The entry is data like any other |
@@ -1370,9 +1602,12 @@ only if it affects the bytes the page is built from:
1370
1602
  | The recovered region (universal mode) disagrees with the same bytes read directly, in the EOCD's two comment-length bytes only | Expected, not an error. A recovered region always declares a zero-length comment (§4.5), so it differs here from any archive written in the declared form (§4.2). Compare the two only up to those bytes |
1371
1603
  | The recovered region (universal mode) disagrees with the same bytes read directly, anywhere else | The file is not well-formed, whichever side is at fault, and a reader that has both MUST NOT silently merge them or pick per entry. Prefer the direct read — it is the writer's own output, where the recovered region is a reconstruction of it — and surface the disagreement rather than displaying either as intact |
1372
1604
 
1373
- Anything the format does not constrain, a reader MUST NOT reject: entries may carry
1374
- any extra fields, timestamps, data descriptors or name-encoding flags a ZIP writer
1375
- would ordinarily emit, and none of it is specified here.
1605
+ Anything the format does not constrain, a reader MUST NOT reject: entries may carry any
1606
+ extra fields, timestamps or data descriptors a ZIP writer would ordinarily emit. The few
1607
+ this document does constrain — the `0x9901` field of an encrypted entry (§4.2), the
1608
+ zip64 records (§5.7), the name-encoding flag (§5.8) — say how an entry is read, not
1609
+ whether it is acceptable, so this row covers them too: each is something a reader meets
1610
+ and reads.
1376
1611
 
1377
1612
  ## 8. Appendices
1378
1613
 
@@ -1445,15 +1680,25 @@ measured except Apple's `ditto`.
1445
1680
  ### 8.2 Anatomy of a small archive
1446
1681
 
1447
1682
  Offsets in `universal.sfz.html` (123077 bytes, two entries, saved from `example.com`
1448
- with the §8.3 command, against core 1.5.108).
1683
+ with the §8.3 command, on the build named there).
1449
1684
  The layout is the *universal* row of the byte map (§3).
1450
1685
 
1686
+ These numbers are one capture, not a contract. Everything from the bootstrap onward
1687
+ moves whenever the inlined ZIP library changes size, so treat the table as an
1688
+ illustration of the shape and not as values to compare a file against. What *is* fixed
1689
+ is the set of relations between the rows — the doctype opening the file with the root
1690
+ element start tag immediately after it, the charset declaration immediately after that
1691
+ and the comment immediately after that, the identifier's twelve bytes ahead of the
1692
+ region, the EOCD's directory offset being an absolute file position, and the entry
1693
+ order. Those are checked by `test/sfz-harness/byte-map.js`, which builds an equivalent
1694
+ specimen without a network.
1695
+
1451
1696
  | Offset | Bytes | Region |
1452
1697
  |---|---|---|
1453
1698
  | 0 | `<!DOCTYPE html>` | `html-prologue` begins |
1454
- | 16 | `<html data-sfz>` | root element start tag; the attribute is the reference implementation's own marker (§1.3) |
1455
- | 31 | `<meta charset=windows-1252>` | the charset rule, inside the first 1024 bytes (§2.1) |
1456
- | 58 | `<!--` … `-->` (ends at 200) | comment written by the implementation, not part of the format; it follows the charset declaration so it cannot push it out of the prescan window |
1699
+ | 15 | `<html data-sfz>` | root element start tag; the attribute is the reference implementation's own marker (§1.3) |
1700
+ | 30 | `<meta charset=windows-1252>` | the charset rule, inside the first 1024 bytes (§2.1) |
1701
+ | 57 | `<!--` … `-->` (ends at 200) | comment written by the implementation; its content is implementation-defined, but where it may appear is not (§3.1, §4.6, §5.6). It follows the charset declaration so it cannot push it out of the prescan window |
1457
1702
  | 200 | `<title>` … `</title>` (ends at 229) | the page title, as numeric character references (§4.6) |
1458
1703
  | 677 | `<style>` | the stylesheet of the blank-page backstop (§4.1) |
1459
1704
  | 855 | `<body hidden>` | start of the blank-page backstop (§4.1) |
@@ -1479,7 +1724,11 @@ Generated with the command-line client running `single-file-core` against
1479
1724
  results of §8.1 were measured on the 1.5.107 build of the same specimen set, which
1480
1725
  differs only inside the prologue and so falls in the same classes: those are grouped
1481
1726
  by whether bytes precede the archive and follow the EOCD, which no prologue change
1482
- alters. `--compress-content` makes the output an archive; `extract-data-from-page`
1727
+ alters. The declared-form results are the exception, `declareAppendedData` being later
1728
+ than that build (§8.5); they were measured separately on a build that has it. The
1729
+ specimen names carry a `.sfz.html` suffix chosen for the harness; the conventions of
1730
+ §2.2 are what the clients produce, not what these files are called.
1731
+ `--compress-content` makes the output an archive; `extract-data-from-page`
1483
1732
  defaults to true there, so the plain variant has to switch it off:
1484
1733
 
1485
1734
  | Specimen | Command |
@@ -1497,20 +1746,18 @@ defaults to true there, so the plain variant has to switch it off:
1497
1746
  These specimens are deliberately small, and a reader tested only against them is
1498
1747
  undertested: they are all flat archives of two or three entries. None
1499
1748
  exercises a root directory, `frames/<n>/` nesting, a second `index.html`, a `data:`-URL
1500
- entry comment, the optional text body or table of contents (§4.6), a UTF-8 BOM, zip64
1749
+ entry comment, the optional text body (§4.6), a UTF-8 BOM, zip64
1501
1750
  (§5.7), a payload past the 64 KB budget, or a relocated reservation with padding left
1502
1751
  in it. Two omissions matter more than the rest, because they are the parts of §5.1 a
1503
1752
  writer is most likely to get wrong: no specimen defeats a rung by its **start**
1504
1753
  pattern, and none defeats one with an **upper-case** pattern. A writer that tested only
1505
- end patterns, or matched them case-sensitively, produces every specimen here unchanged
1506
- — and the first of those two mistakes is one the reference writer actually shipped
1507
- (§8.5).
1754
+ end patterns, or matched them case-sensitively, produces every specimen here unchanged.
1508
1755
 
1509
1756
  Two specimens cannot be produced from a URL alone. The **ladder** specimen, which
1510
1757
  forces the second rung of §5.1, needs a page referencing an image whose stored bytes
1511
1758
  contain `-->`; the archive then wraps in `<script type=sfz-data>`. The **zip64** specimen requires
1512
1759
  an archive past the thresholds of §5.7, so it is produced by calling the writer
1513
- directly with zip64 forced on the ZIP writer.
1760
+ directly with zip64 forced on the ZIP writer, as `test/sfz-harness/zip64.js` does.
1514
1761
 
1515
1762
  The measurements quoted elsewhere in this document come from the same harness: the
1516
1763
  payload growth rate of §5.2 (86 KB → 181 bytes, 283 KB → 465, 1.07 MB → 1645, 4.2 MB
@@ -1535,7 +1782,7 @@ multi-byte ones (`utf-8`, `utf-16le`, `utf-16be`, `gbk`, `gb18030`, `big5`, `euc
1535
1782
  `shift_jis`, `euc-kr`, `iso-2022-jp`) decode a lone byte sequence to U+FFFD or to fewer
1536
1783
  than 256 characters.
1537
1784
  The reverse table each one needs ranges from 8 entries (`iso-8859-15`) to 128
1538
- (`koi8-r`, `koi8-u` and `ibm866`); windows-1252 needs 27.
1785
+ (`koi8-r`, `koi8-u`, `ibm866` and `x-user-defined`); windows-1252 needs 27.
1539
1786
 
1540
1787
  **That the round trip is charset-independent.** The mechanism of §5.5 — parse, then
1541
1788
  re-encode with the reverse table, restoring newlines from the 2-bit codes and NUL from
@@ -1562,11 +1809,13 @@ predicts.
1562
1809
  | August 2026 | Core 1.5.108: the ZIP region carries the identifier `sfz-data` and the extractor addresses it with that instead of deducing it from its position beside `<sfz-extra-data>` (§4.5). This fixes universal extraction on the `<style type=sfz-data>` rung, where the reference extractor's own relocation of `style` elements into the head moved the region out from under the positional rule |
1563
1810
  | August 2026 | Core 1.5.108: the recovery payload stops two bytes short of the End Of Central Directory record, excluding its comment-length field (§1.3), which lets universal-mode archives declare their appended data as the archive comment — a writer option, for `java.util.zip` and the readers that reject undeclared trailing bytes (§4.2) |
1564
1811
  | August 2026 | Core 1.5.108: the PDF and PNG faces test a wrapper rung's start pattern as well as its end pattern, closing the same script-data escape hole the ZIP region was already guarded against — a face payload holding `<!--` and then `<script` took the `<script type=sfz-data>` rung and swallowed the rest of the document (§5.1) |
1565
- | August 2026 | Core 1.5.108: the retry loop discards a relocation reservation at most once per build, so a payload sitting on the appended-data boundary cannot oscillate between the two placements forever (§6.2) |
1812
+ | August 2026 | Core 1.5.108: the retry loop never discards a relocation reservation, so a payload sitting on the appended-data boundary cannot oscillate between the two placements forever (§6.2) |
1566
1813
  | August 2026 | Core 1.5.110: a PDF or PNG face whose payload names every rung is dropped instead of written bare (§5.1). Found by nesting an archive inside itself as both faces: the fifth level exhausts the ladder, and readers then extracted the fourth level's archive — checksums intact, no way to tell (§7.4) |
1567
1814
  | August 2026 | Core 1.5.110: a PNG face leaving the comment rung on its checksum resumes the rung search instead of taking the next rung untested (§5.1). Taking it put a payload holding `</script>` on the script rung, where its own bytes closed the wrapper 93 bytes in and left the image data, the chunk framing and the whole ZIP region to the parser |
1568
1815
  | August 2026 | Core 1.5.110: `<svg><![CDATA[` joins the ladder above `<plaintext>` (§5.1) — the one rung whose terminator, `]]>`, real payloads rarely carry. It gives a payload naming every element rung somewhere to go that does not cost the appended-data placement, and moves the self-nesting limit from the fifth level to the sixth |
1569
1816
  | August 2026 | Core 1.5.115: password-protected archives withhold the provenance comment and the canonical link as well (§5.6). Both wrote the page's own URL into the prologue, beside the title that was already withheld, so the address the archive was saved from stayed in the clear |
1817
+ | August 2026 | Core 1.5.119: the inlined ZIP library is built ASCII-only, and §2.1 now requires it of any bootstrap. Its CP437 table had been emitted as literal characters, which the page re-decoded as windows-1252, growing the table from 256 entries to 508 and shifting every lookup by 60 — so the one entry read without the UTF-8 flag, `page.pdf`, came back mangled and no archive with a PDF face extracted in any engine (§5.8) |
1818
+ | August 2026 | Core 1.5.120: the hand-built `page.pdf` records set the language encoding flag, like every entry the ZIP writer produces (§5.8). Its name is ASCII, so no decoded name changes; what changes is that no entry in an archive is read through CP437 any more, closing the path the 1.5.119 defect surfaced on |
1570
1819
 
1571
1820
  This document was itself revised in August 2026, against core 1.5.108, after several
1572
1821
  independent reviews. One of them was a reader built from this specification alone, with
@@ -1603,3 +1852,13 @@ BOM, a user override and a transport-layer charset all outranking it — narrow
1603
1852
  practice, since the raw read comes first and no encoding applies to it (§2.1) — and that
1604
1853
  the recovery payload's 32-bit length field caps the region below 2^32 bytes, with
1605
1854
  engine string limits binding well before that (§5.5).
1855
+
1856
+ A pass in September 2026, against core 1.5.120, read the text alone first and then
1857
+ checked each open question against the writer. It corrected two statements about the
1858
+ reference writer that the code contradicted: the retry loop never discards a
1859
+ reservation, and the archive time is read once per build, not once per pass (§6.2).
1860
+ It added what only the code could say: the PNG build steps that §6.1 had skipped, the
1861
+ fields patched after the wrapper check and the size below which they are harmless
1862
+ (§5.1), the chunks the PNG face copies (§3.1), the per-frame manifests and the root
1863
+ directory's name (§7.1), the `page.pdf` header fields (§6.1), and the range-reading
1864
+ failure path (§4.1).