@clearmist-labs/comic-archive-handler 1.2.0 → 1.4.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/docs/API.md CHANGED
@@ -27,7 +27,7 @@ console.log(type); // 'zip'
27
27
 
28
28
  ### `isZip(input)`, `isRar(input)`, `isTar(input)`, `isAsar(input)`, `is7z(input)`, `isAce(input)`
29
29
 
30
- Convenience predicates that detect an archive and return whether it matches the named format. `isAce` detects ACE (`.cba`) archives, but every actual ACE operation (read, write, or conversion) throws `UnsupportedOperationError` — see [`UnsupportedOperationError`](#unsupportedoperationerror).
30
+ Convenience predicates that detect an archive and return whether it matches the named format. `isAce` detects ACE (`.cba`) archives, but every actual ACE operation (read, write, or conversion) throws `UnsupportedOperationError`. See [`UnsupportedOperationError`](#unsupportedoperationerror).
31
31
 
32
32
  **Options**
33
33
 
@@ -102,7 +102,7 @@ const page = await cah.readArchiveEntry(comicBuffer, 'P00001.jpg');
102
102
 
103
103
  ### `readArchiveEntries(input)`
104
104
 
105
- Reads every entry's contents and SHA256 in a single pass over the archive, yielding `{ path, size?, buffer, sha256 }` in archive iteration order. Prefer this over calling `readArchiveEntry`/`sha256ArchiveEntry` once per entry: zip and asar support real random access, but the sequential/CLI-driven formats (rar, 7z, ace, tar) re-scan the archive from the start on every such call — reading every entry that way costs O(n^2) instead of O(n) for those formats. It also avoids reading each entry's data twice (once for the hash, once for the buffer).
105
+ Reads every entry's contents and SHA256 in a single pass over the archive, yielding `{ path, size?, buffer, sha256 }` in archive iteration order. Prefer this over calling `readArchiveEntry`/`sha256ArchiveEntry` once per entry: zip and asar support real random access, but the sequential/CLI-driven formats (rar, 7z, ace, tar) re-scan the archive from the start on every such call: reading every entry that way costs O(n^2) instead of O(n) for those formats. It also avoids reading each entry's data twice (once for the hash, once for the buffer).
106
106
 
107
107
  **Options**
108
108
 
@@ -172,9 +172,26 @@ const cleaned = await cah.stripNonEssentialFiles(comicBuffer, {
172
172
  });
173
173
  ```
174
174
 
175
+ ### `removeArchiveEntry(input, entryPath, options?)`
176
+
177
+ Removes a single named entry from an archive, leaving every other entry unchanged. Throws `MetadataNotFoundError` when no entry matches `entryPath`.
178
+
179
+ **Options**
180
+
181
+ - `input: ArchiveInput` - A filesystem path or archive `Buffer`.
182
+ - `entryPath: string` - Exact archive entry path to remove.
183
+ - `options.tempDir?: string` - Temporary staging directory for ASAR or 7z operations.
184
+ - `options.output?: string | Writable` - Output destination. Without it, returns a `Buffer`.
185
+
186
+ **Example**
187
+
188
+ ```js
189
+ const trimmed = await cah.removeArchiveEntry(comicBuffer, 'thumbs.db');
190
+ ```
191
+
175
192
  ## Metadata
176
193
 
177
- `ComicMetadata` is the canonical metadata shape, covering every field defined by both bundled schemas (`schemas/ComicInfo v2.1.xsd` and `schemas/MetronInfo v1.1.xsd`) — see [`schemaVersions.ts`](../src/metadata/schemaVersions.ts) if those XSDs are ever upgraded. Conversion between ComicInfo.xml and MetronInfo.xml is intentionally lossy when a field has no equivalent in the target schema; the complete field mapping is maintained in `src/metadata/schema.ts`.
194
+ `ComicMetadata` is the canonical metadata shape, covering every field defined by both bundled schemas (`schemas/ComicInfo v2.1.xsd` and `schemas/MetronInfo v1.1.xsd`). See [`schemaVersions.ts`](../src/metadata/schemaVersions.ts) if those XSDs are ever upgraded. Conversion between ComicInfo.xml and MetronInfo.xml is intentionally lossy when a field has no equivalent in the target schema; the complete field mapping is maintained in `src/metadata/schema.ts`.
178
195
 
179
196
  List fields that MetronInfo represents as an element with an optional `id` attribute (`genres`, `tags`, `characters`, `teams`, `locations`, `stories`, `reprints`) accept either a plain string or a `{ name, id? }` object (see [`MetronResource`](#types)); the `id` is preserved on round-trip through MetronInfo.xml and dropped when writing ComicInfo.xml, which has no equivalent concept.
180
197
 
@@ -266,7 +283,7 @@ const metadata = cah.xmlToMetadata(xml, 'ComicInfo');
266
283
 
267
284
  ### `validateMetadataXml(xml, schema)`
268
285
 
269
- Validates an XML document against the bundled XSD for `'ComicInfo'` or `'MetronInfo'`, using [`xmllint-wasm`](https://www.npmjs.com/package/xmllint-wasm) (a WASM build of libxml2 — no native or Java dependency). Returns `{ valid, issues }`; `issues` is empty when the document conforms, otherwise it has one entry per schema violation with the offending element/attribute in `message` and, where available, a `line` number.
286
+ Validates an XML document against the bundled XSD for `'ComicInfo'` or `'MetronInfo'`, using [`xmllint-wasm`](https://www.npmjs.com/package/xmllint-wasm) (a WASM build of libxml2; no native or Java dependency). Returns `{ valid, issues }`; `issues` is empty when the document conforms, otherwise it has one entry per schema violation with the offending element/attribute in `message` and, where available, a `line` number.
270
287
 
271
288
  Because the validator implements XSD 1.0, it cannot check the two `<xs:assert>` business rules in MetronInfo.xml v1.1 (at most one primary `URL`, at most one primary `ID`); everything else in both schemas is enforced.
272
289
 
@@ -289,9 +306,21 @@ if (!valid) {
289
306
  }
290
307
  ```
291
308
 
309
+ **ASAR is the one exception to everything below**: an asar archive never
310
+ stores metadata as a `ComicInfo.xml`/`MetronInfo.xml` entry. Instead,
311
+ `hasComicMetadata`/`readArchiveMetadata`/`addMetadataToArchive`/
312
+ `removeComicMetadata` all read/write a `comicMetadata` key embedded directly
313
+ in the asar container's own header (the small JSON block at the start of
314
+ every `.asar` file, alongside the header's existing `files` key), as plain
315
+ JSON; no XML involved. This is transparent to callers: the same four
316
+ functions work identically regardless of archive type, and `schema` still
317
+ selects `'ComicInfo'` vs `'MetronInfo'` field semantics either way. A
318
+ `ComicInfo.xml`/`MetronInfo.xml` _entry_ inside an asar archive is never
319
+ read as metadata by any of these functions.
320
+
292
321
  ### `hasComicMetadata(input)`
293
322
 
294
- Checks for a root-level or nested `ComicInfo.xml` or `MetronInfo.xml` entry without parsing it. Returns `{ present, schema?, path? }`.
323
+ Checks for a root-level or nested `ComicInfo.xml` or `MetronInfo.xml` entry without parsing it (or, for asar, a `comicMetadata` header key). Returns `{ present, schema?, path?, bothPresent? }`. `path` is only set for a real archive entry (never for asar); `bothPresent` (asar only) is `true` when both `ComicInfo` and `MetronInfo` are present in the header.
295
324
 
296
325
  **Options**
297
326
 
@@ -306,13 +335,14 @@ if (result.present) {
306
335
  }
307
336
  ```
308
337
 
309
- ### `readArchiveMetadata(input)`
338
+ ### `readArchiveMetadata(input, schema?)`
310
339
 
311
- Finds and parses the first recognized metadata entry. Returns `{ schema, metadata }`, or `null` when no metadata file is present.
340
+ Finds and parses metadata. With no `schema`, returns whichever is found first (asar checks ComicInfo before MetronInfo). Pass `schema` to read that one specifically regardless of preference. The only way to read a non-preferred schema back out of an asar archive that has both, since asar metadata isn't addressable by path. Returns `{ schema, metadata }`, or `null` when that metadata isn't present.
312
341
 
313
342
  **Options**
314
343
 
315
344
  - `input: ArchiveInput` - A filesystem path or archive `Buffer`.
345
+ - `schema?: 'ComicInfo' | 'MetronInfo'` - Read this specific schema instead of whichever is preferred.
316
346
 
317
347
  **Example**
318
348
 
@@ -323,7 +353,7 @@ console.log(result?.metadata.title);
323
353
 
324
354
  ### `addMetadataToArchive(input, metadata, schema, options?)`
325
355
 
326
- Adds a `ComicInfo.xml` or `MetronInfo.xml` entry to an archive. Existing metadata is rejected unless `overwrite: true` is supplied.
356
+ Adds a `ComicInfo.xml` or `MetronInfo.xml` entry to an archive (or, for asar, sets the `comicMetadata.{schema}` header key). Existing metadata is rejected unless `overwrite: true` is supplied.
327
357
 
328
358
  **Options**
329
359
 
@@ -349,6 +379,23 @@ const withMetadata = await cah.addMetadataToArchive(
349
379
  );
350
380
  ```
351
381
 
382
+ ### `removeComicMetadata(input, schema, options?)`
383
+
384
+ Removes the embedded `schema` metadata (a `ComicInfo.xml`/`MetronInfo.xml` entry, or, for asar, the `comicMetadata.{schema}` header key). Throws `ArchiveFormatError` if that schema isn't present. Check `hasComicMetadata` first if you want a no-op instead.
385
+
386
+ **Options**
387
+
388
+ - `input: ArchiveInput` - A filesystem path or archive `Buffer`.
389
+ - `schema: 'ComicInfo' | 'MetronInfo'` - Metadata to remove.
390
+ - `options.tempDir?: string` - Temporary staging directory for ASAR or 7z operations.
391
+ - `options.output?: string | Writable` - Output destination. Without it, returns a `Buffer`.
392
+
393
+ **Example**
394
+
395
+ ```js
396
+ const withoutMetronInfo = await cah.removeComicMetadata(comicBuffer, 'MetronInfo');
397
+ ```
398
+
352
399
  ## Image conversion and detection
353
400
 
354
401
  ### `convertImageBuffer(image, format, options?)`
@@ -535,7 +582,7 @@ const digest = await cah.sha256Archive('/books/example.cbz');
535
582
 
536
583
  ### `benchmarkArchive(filePath, options?)`
537
584
 
538
- Extracts a comic archive into the report's `source/` subdirectory, validates it contains image files, then generates one archive per writable container format (`zip`, `tar`, `asar`, `7z`) times benchmarked page image format (`webp`, `png`, `jpg` by default, or a subset via `options.imageFormats`) — 12 variants in total by default, each containing only the original's XML and image entries, always re-encoded from the untouched source pages (never from a previously-converted variant). Benchmarks each variant's average archive-creation time (packaging only — images are re-encoded once per format before timing starts) and average random single-entry read time, writes the generated archives and a markdown report to a timestamped subdirectory, and returns the results.
585
+ Extracts a comic archive into the report's `source/` subdirectory, validates it contains image files, then generates one archive per writable container format (`zip`, `tar`, `asar`, `7z`) times benchmarked page image format (`webp`, `png`, `jpg` by default, or a subset via `options.imageFormats`). 12 variants in total by default, each containing only the original's XML and image entries, always re-encoded from the untouched source pages (never from a previously-converted variant). Benchmarks each variant's average archive-creation time (packaging only; images are re-encoded once per format before timing starts) and average random single-entry read time, writes the generated archives and a markdown report to a timestamped subdirectory, and returns the results.
539
586
 
540
587
  Throws `ArchiveFormatError` (undetectable format) or `UnsupportedOperationError` (ACE) if `filePath` isn't extractable, `NoImagesFoundError` if the archive has no image files, `FilesystemAccessError` if `filePath` doesn't exist, and `RangeError` if `options.imageFormats` is empty or names an unsupported format.
541
588
 
@@ -605,6 +652,7 @@ The following types are exported for TypeScript consumers.
605
652
  - `AddMetadataOptions` - Archive write options plus `overwrite?: boolean`.
606
653
  - `RenameOptions` - Archive write options plus `start?: number` and `pad?: number`.
607
654
  - `StripOptions` - Archive write options plus `extraKeepExtensions?: string[]`.
655
+ - `RemoveEntryOptions` - Archive write options (no additional fields).
608
656
  - `ImageConvertOptions` - `webp?`, `jpeg?`, and `png?` format option groups.
609
657
  - `WebpOptions` - `quality?`, `effort?`, and `smartSubsample?`.
610
658
  - `JpegOptions` - `quality?`.
@@ -704,7 +752,7 @@ console.log(cah.joinResourceNames([{ name: 'Action', id: 'genre-1' }, 'Adventure
704
752
 
705
753
  ### Schema reference enums
706
754
 
707
- Readonly value lists for the enumerated (`xs:enumeration`) types in each XSD. These are reference data, not enforced constraints — fields such as `ageRating` remain plain `string` so unrecognized or future values still round-trip.
755
+ Readonly value lists for the enumerated (`xs:enumeration`) types in each XSD. These are reference data, not enforced constraints. Fields such as `ageRating` remain plain `string` so unrecognized or future values still round-trip.
708
756
 
709
757
  - `COMIC_INFO_YES_NO_VALUES` - `'Unknown' | 'No' | 'Yes'` (ComicInfo `BlackAndWhite`).
710
758
  - `COMIC_INFO_MANGA_VALUES` - `'Unknown' | 'No' | 'Yes' | 'YesAndRightToLeft'` (ComicInfo `Manga`).
@@ -713,7 +761,7 @@ Readonly value lists for the enumerated (`xs:enumeration`) types in each XSD. Th
713
761
  - `METRON_FORMAT_VALUES` - MetronInfo `Series > Format` values, e.g. `'Single Issue'`, `'Trade Paperback'`.
714
762
  - `METRON_INFORMATION_SOURCE_VALUES` - MetronInfo `IDS > ID` `source` attribute values, e.g. `'Comic Vine'`, `'Metron'`.
715
763
  - `METRON_ROLE_VALUES` - MetronInfo `Credits > Credit > Roles > Role` values (a larger set than `COMIC_INFO_CREDIT_ROLES`).
716
- - `METRON_AGE_RATING_VALUES` - MetronInfo `AgeRating` values — a different set from ComicInfo's.
764
+ - `METRON_AGE_RATING_VALUES` - MetronInfo `AgeRating` values; a different set from ComicInfo's.
717
765
 
718
766
  **Example**
719
767
 
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@clearmist-labs/comic-archive-handler",
3
- "version": "1.2.0",
3
+ "version": "1.4.0",
4
4
  "description": "Detect, convert, and manipulate comic book archives: container conversion, edit metadata, image re-encoding, perceptual + content hashing.",
5
5
  "keywords": [
6
6
  "7z",