@clearmist-labs/comic-archive-handler 1.3.0 → 1.4.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/docs/API.md CHANGED
@@ -27,7 +27,7 @@ console.log(type); // 'zip'
27
27
 
28
28
  ### `isZip(input)`, `isRar(input)`, `isTar(input)`, `isAsar(input)`, `is7z(input)`, `isAce(input)`
29
29
 
30
- Convenience predicates that detect an archive and return whether it matches the named format. `isAce` detects ACE (`.cba`) archives, but every actual ACE operation (read, write, or conversion) throws `UnsupportedOperationError` — see [`UnsupportedOperationError`](#unsupportedoperationerror).
30
+ Convenience predicates that detect an archive and return whether it matches the named format. `isAce` detects ACE (`.cba`) archives, but every actual ACE operation (read, write, or conversion) throws `UnsupportedOperationError`. See [`UnsupportedOperationError`](#unsupportedoperationerror).
31
31
 
32
32
  **Options**
33
33
 
@@ -102,7 +102,7 @@ const page = await cah.readArchiveEntry(comicBuffer, 'P00001.jpg');
102
102
 
103
103
  ### `readArchiveEntries(input)`
104
104
 
105
- Reads every entry's contents and SHA256 in a single pass over the archive, yielding `{ path, size?, buffer, sha256 }` in archive iteration order. Prefer this over calling `readArchiveEntry`/`sha256ArchiveEntry` once per entry: zip and asar support real random access, but the sequential/CLI-driven formats (rar, 7z, ace, tar) re-scan the archive from the start on every such call — reading every entry that way costs O(n^2) instead of O(n) for those formats. It also avoids reading each entry's data twice (once for the hash, once for the buffer).
105
+ Reads every entry's contents and SHA256 in a single pass over the archive, yielding `{ path, size?, buffer, sha256 }` in archive iteration order. Prefer this over calling `readArchiveEntry`/`sha256ArchiveEntry` once per entry: zip and asar support real random access, but the sequential/CLI-driven formats (rar, 7z, ace, tar) re-scan the archive from the start on every such call: reading every entry that way costs O(n^2) instead of O(n) for those formats. It also avoids reading each entry's data twice (once for the hash, once for the buffer).
106
106
 
107
107
  **Options**
108
108
 
@@ -191,7 +191,7 @@ const trimmed = await cah.removeArchiveEntry(comicBuffer, 'thumbs.db');
191
191
 
192
192
  ## Metadata
193
193
 
194
- `ComicMetadata` is the canonical metadata shape, covering every field defined by both bundled schemas (`schemas/ComicInfo v2.1.xsd` and `schemas/MetronInfo v1.1.xsd`) — see [`schemaVersions.ts`](../src/metadata/schemaVersions.ts) if those XSDs are ever upgraded. Conversion between ComicInfo.xml and MetronInfo.xml is intentionally lossy when a field has no equivalent in the target schema; the complete field mapping is maintained in `src/metadata/schema.ts`.
194
+ `ComicMetadata` is the canonical metadata shape, covering every field defined by both bundled schemas (`schemas/ComicInfo v2.1.xsd` and `schemas/MetronInfo v1.1.xsd`). See [`schemaVersions.ts`](../src/metadata/schemaVersions.ts) if those XSDs are ever upgraded. Conversion between ComicInfo.xml and MetronInfo.xml is intentionally lossy when a field has no equivalent in the target schema; the complete field mapping is maintained in `src/metadata/schema.ts`.
195
195
 
196
196
  List fields that MetronInfo represents as an element with an optional `id` attribute (`genres`, `tags`, `characters`, `teams`, `locations`, `stories`, `reprints`) accept either a plain string or a `{ name, id? }` object (see [`MetronResource`](#types)); the `id` is preserved on round-trip through MetronInfo.xml and dropped when writing ComicInfo.xml, which has no equivalent concept.
197
197
 
@@ -283,7 +283,7 @@ const metadata = cah.xmlToMetadata(xml, 'ComicInfo');
283
283
 
284
284
  ### `validateMetadataXml(xml, schema)`
285
285
 
286
- Validates an XML document against the bundled XSD for `'ComicInfo'` or `'MetronInfo'`, using [`xmllint-wasm`](https://www.npmjs.com/package/xmllint-wasm) (a WASM build of libxml2 — no native or Java dependency). Returns `{ valid, issues }`; `issues` is empty when the document conforms, otherwise it has one entry per schema violation with the offending element/attribute in `message` and, where available, a `line` number.
286
+ Validates an XML document against the bundled XSD for `'ComicInfo'` or `'MetronInfo'`, using [`xmllint-wasm`](https://www.npmjs.com/package/xmllint-wasm) (a WASM build of libxml2; no native or Java dependency). Returns `{ valid, issues }`; `issues` is empty when the document conforms, otherwise it has one entry per schema violation with the offending element/attribute in `message` and, where available, a `line` number.
287
287
 
288
288
  Because the validator implements XSD 1.0, it cannot check the two `<xs:assert>` business rules in MetronInfo.xml v1.1 (at most one primary `URL`, at most one primary `ID`); everything else in both schemas is enforced.
289
289
 
@@ -306,9 +306,21 @@ if (!valid) {
306
306
  }
307
307
  ```
308
308
 
309
+ **ASAR is the one exception to everything below**: an asar archive never
310
+ stores metadata as a `ComicInfo.xml`/`MetronInfo.xml` entry. Instead,
311
+ `hasComicMetadata`/`readArchiveMetadata`/`addMetadataToArchive`/
312
+ `removeComicMetadata` all read/write a `comicMetadata` key embedded directly
313
+ in the asar container's own header (the small JSON block at the start of
314
+ every `.asar` file, alongside the header's existing `files` key), as plain
315
+ JSON; no XML involved. This is transparent to callers: the same four
316
+ functions work identically regardless of archive type, and `schema` still
317
+ selects `'ComicInfo'` vs `'MetronInfo'` field semantics either way. A
318
+ `ComicInfo.xml`/`MetronInfo.xml` _entry_ inside an asar archive is never
319
+ read as metadata by any of these functions.
320
+
309
321
  ### `hasComicMetadata(input)`
310
322
 
311
- Checks for a root-level or nested `ComicInfo.xml` or `MetronInfo.xml` entry without parsing it. Returns `{ present, schema?, path? }`.
323
+ Checks for a root-level or nested `ComicInfo.xml` or `MetronInfo.xml` entry without parsing it (or, for asar, a `comicMetadata` header key). Returns `{ present, schema?, path?, bothPresent? }`. `path` is only set for a real archive entry (never for asar); `bothPresent` (asar only) is `true` when both `ComicInfo` and `MetronInfo` are present in the header.
312
324
 
313
325
  **Options**
314
326
 
@@ -323,13 +335,14 @@ if (result.present) {
323
335
  }
324
336
  ```
325
337
 
326
- ### `readArchiveMetadata(input)`
338
+ ### `readArchiveMetadata(input, schema?)`
327
339
 
328
- Finds and parses the first recognized metadata entry. Returns `{ schema, metadata }`, or `null` when no metadata file is present.
340
+ Finds and parses metadata. With no `schema`, returns whichever is found first (asar checks ComicInfo before MetronInfo). Pass `schema` to read that one specifically regardless of preference. The only way to read a non-preferred schema back out of an asar archive that has both, since asar metadata isn't addressable by path. Returns `{ schema, metadata }`, or `null` when that metadata isn't present.
329
341
 
330
342
  **Options**
331
343
 
332
344
  - `input: ArchiveInput` - A filesystem path or archive `Buffer`.
345
+ - `schema?: 'ComicInfo' | 'MetronInfo'` - Read this specific schema instead of whichever is preferred.
333
346
 
334
347
  **Example**
335
348
 
@@ -340,7 +353,7 @@ console.log(result?.metadata.title);
340
353
 
341
354
  ### `addMetadataToArchive(input, metadata, schema, options?)`
342
355
 
343
- Adds a `ComicInfo.xml` or `MetronInfo.xml` entry to an archive. Existing metadata is rejected unless `overwrite: true` is supplied.
356
+ Adds a `ComicInfo.xml` or `MetronInfo.xml` entry to an archive (or, for asar, sets the `comicMetadata.{schema}` header key). Existing metadata is rejected unless `overwrite: true` is supplied.
344
357
 
345
358
  **Options**
346
359
 
@@ -366,6 +379,23 @@ const withMetadata = await cah.addMetadataToArchive(
366
379
  );
367
380
  ```
368
381
 
382
+ ### `removeComicMetadata(input, schema, options?)`
383
+
384
+ Removes the embedded `schema` metadata (a `ComicInfo.xml`/`MetronInfo.xml` entry, or, for asar, the `comicMetadata.{schema}` header key). Throws `ArchiveFormatError` if that schema isn't present. Check `hasComicMetadata` first if you want a no-op instead.
385
+
386
+ **Options**
387
+
388
+ - `input: ArchiveInput` - A filesystem path or archive `Buffer`.
389
+ - `schema: 'ComicInfo' | 'MetronInfo'` - Metadata to remove.
390
+ - `options.tempDir?: string` - Temporary staging directory for ASAR or 7z operations.
391
+ - `options.output?: string | Writable` - Output destination. Without it, returns a `Buffer`.
392
+
393
+ **Example**
394
+
395
+ ```js
396
+ const withoutMetronInfo = await cah.removeComicMetadata(comicBuffer, 'MetronInfo');
397
+ ```
398
+
369
399
  ## Image conversion and detection
370
400
 
371
401
  ### `convertImageBuffer(image, format, options?)`
@@ -552,7 +582,7 @@ const digest = await cah.sha256Archive('/books/example.cbz');
552
582
 
553
583
  ### `benchmarkArchive(filePath, options?)`
554
584
 
555
- Extracts a comic archive into the report's `source/` subdirectory, validates it contains image files, then generates one archive per writable container format (`zip`, `tar`, `asar`, `7z`) times benchmarked page image format (`webp`, `png`, `jpg` by default, or a subset via `options.imageFormats`) — 12 variants in total by default, each containing only the original's XML and image entries, always re-encoded from the untouched source pages (never from a previously-converted variant). Benchmarks each variant's average archive-creation time (packaging only — images are re-encoded once per format before timing starts) and average random single-entry read time, writes the generated archives and a markdown report to a timestamped subdirectory, and returns the results.
585
+ Extracts a comic archive into the report's `source/` subdirectory, validates it contains image files, then generates one archive per writable container format (`zip`, `tar`, `asar`, `7z`) times benchmarked page image format (`webp`, `png`, `jpg` by default, or a subset via `options.imageFormats`). 12 variants in total by default, each containing only the original's XML and image entries, always re-encoded from the untouched source pages (never from a previously-converted variant). Benchmarks each variant's average archive-creation time (packaging only; images are re-encoded once per format before timing starts) and average random single-entry read time, writes the generated archives and a markdown report to a timestamped subdirectory, and returns the results.
556
586
 
557
587
  Throws `ArchiveFormatError` (undetectable format) or `UnsupportedOperationError` (ACE) if `filePath` isn't extractable, `NoImagesFoundError` if the archive has no image files, `FilesystemAccessError` if `filePath` doesn't exist, and `RangeError` if `options.imageFormats` is empty or names an unsupported format.
558
588
 
@@ -722,7 +752,7 @@ console.log(cah.joinResourceNames([{ name: 'Action', id: 'genre-1' }, 'Adventure
722
752
 
723
753
  ### Schema reference enums
724
754
 
725
- Readonly value lists for the enumerated (`xs:enumeration`) types in each XSD. These are reference data, not enforced constraints — fields such as `ageRating` remain plain `string` so unrecognized or future values still round-trip.
755
+ Readonly value lists for the enumerated (`xs:enumeration`) types in each XSD. These are reference data, not enforced constraints. Fields such as `ageRating` remain plain `string` so unrecognized or future values still round-trip.
726
756
 
727
757
  - `COMIC_INFO_YES_NO_VALUES` - `'Unknown' | 'No' | 'Yes'` (ComicInfo `BlackAndWhite`).
728
758
  - `COMIC_INFO_MANGA_VALUES` - `'Unknown' | 'No' | 'Yes' | 'YesAndRightToLeft'` (ComicInfo `Manga`).
@@ -731,7 +761,7 @@ Readonly value lists for the enumerated (`xs:enumeration`) types in each XSD. Th
731
761
  - `METRON_FORMAT_VALUES` - MetronInfo `Series > Format` values, e.g. `'Single Issue'`, `'Trade Paperback'`.
732
762
  - `METRON_INFORMATION_SOURCE_VALUES` - MetronInfo `IDS > ID` `source` attribute values, e.g. `'Comic Vine'`, `'Metron'`.
733
763
  - `METRON_ROLE_VALUES` - MetronInfo `Credits > Credit > Roles > Role` values (a larger set than `COMIC_INFO_CREDIT_ROLES`).
734
- - `METRON_AGE_RATING_VALUES` - MetronInfo `AgeRating` values — a different set from ComicInfo's.
764
+ - `METRON_AGE_RATING_VALUES` - MetronInfo `AgeRating` values; a different set from ComicInfo's.
735
765
 
736
766
  **Example**
737
767
 
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@clearmist-labs/comic-archive-handler",
3
- "version": "1.3.0",
3
+ "version": "1.4.0",
4
4
  "description": "Detect, convert, and manipulate comic book archives: container conversion, edit metadata, image re-encoding, perceptual + content hashing.",
5
5
  "keywords": [
6
6
  "7z",