@clearmist-labs/comic-archive-handler 1.2.0 → 1.4.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +2 -1
- package/dist/index.d.mts +38 -16
- package/dist/index.d.mts.map +1 -1
- package/dist/index.mjs +200 -73
- package/dist/index.mjs.map +1 -1
- package/docs/API.md +59 -11
- package/package.json +1 -1
package/docs/API.md
CHANGED
|
@@ -27,7 +27,7 @@ console.log(type); // 'zip'
|
|
|
27
27
|
|
|
28
28
|
### `isZip(input)`, `isRar(input)`, `isTar(input)`, `isAsar(input)`, `is7z(input)`, `isAce(input)`
|
|
29
29
|
|
|
30
|
-
Convenience predicates that detect an archive and return whether it matches the named format. `isAce` detects ACE (`.cba`) archives, but every actual ACE operation (read, write, or conversion) throws `UnsupportedOperationError
|
|
30
|
+
Convenience predicates that detect an archive and return whether it matches the named format. `isAce` detects ACE (`.cba`) archives, but every actual ACE operation (read, write, or conversion) throws `UnsupportedOperationError`. See [`UnsupportedOperationError`](#unsupportedoperationerror).
|
|
31
31
|
|
|
32
32
|
**Options**
|
|
33
33
|
|
|
@@ -102,7 +102,7 @@ const page = await cah.readArchiveEntry(comicBuffer, 'P00001.jpg');
|
|
|
102
102
|
|
|
103
103
|
### `readArchiveEntries(input)`
|
|
104
104
|
|
|
105
|
-
Reads every entry's contents and SHA256 in a single pass over the archive, yielding `{ path, size?, buffer, sha256 }` in archive iteration order. Prefer this over calling `readArchiveEntry`/`sha256ArchiveEntry` once per entry: zip and asar support real random access, but the sequential/CLI-driven formats (rar, 7z, ace, tar) re-scan the archive from the start on every such call
|
|
105
|
+
Reads every entry's contents and SHA256 in a single pass over the archive, yielding `{ path, size?, buffer, sha256 }` in archive iteration order. Prefer this over calling `readArchiveEntry`/`sha256ArchiveEntry` once per entry: zip and asar support real random access, but the sequential/CLI-driven formats (rar, 7z, ace, tar) re-scan the archive from the start on every such call: reading every entry that way costs O(n^2) instead of O(n) for those formats. It also avoids reading each entry's data twice (once for the hash, once for the buffer).
|
|
106
106
|
|
|
107
107
|
**Options**
|
|
108
108
|
|
|
@@ -172,9 +172,26 @@ const cleaned = await cah.stripNonEssentialFiles(comicBuffer, {
|
|
|
172
172
|
});
|
|
173
173
|
```
|
|
174
174
|
|
|
175
|
+
### `removeArchiveEntry(input, entryPath, options?)`
|
|
176
|
+
|
|
177
|
+
Removes a single named entry from an archive, leaving every other entry unchanged. Throws `MetadataNotFoundError` when no entry matches `entryPath`.
|
|
178
|
+
|
|
179
|
+
**Options**
|
|
180
|
+
|
|
181
|
+
- `input: ArchiveInput` - A filesystem path or archive `Buffer`.
|
|
182
|
+
- `entryPath: string` - Exact archive entry path to remove.
|
|
183
|
+
- `options.tempDir?: string` - Temporary staging directory for ASAR or 7z operations.
|
|
184
|
+
- `options.output?: string | Writable` - Output destination. Without it, returns a `Buffer`.
|
|
185
|
+
|
|
186
|
+
**Example**
|
|
187
|
+
|
|
188
|
+
```js
|
|
189
|
+
const trimmed = await cah.removeArchiveEntry(comicBuffer, 'thumbs.db');
|
|
190
|
+
```
|
|
191
|
+
|
|
175
192
|
## Metadata
|
|
176
193
|
|
|
177
|
-
`ComicMetadata` is the canonical metadata shape, covering every field defined by both bundled schemas (`schemas/ComicInfo v2.1.xsd` and `schemas/MetronInfo v1.1.xsd`)
|
|
194
|
+
`ComicMetadata` is the canonical metadata shape, covering every field defined by both bundled schemas (`schemas/ComicInfo v2.1.xsd` and `schemas/MetronInfo v1.1.xsd`). See [`schemaVersions.ts`](../src/metadata/schemaVersions.ts) if those XSDs are ever upgraded. Conversion between ComicInfo.xml and MetronInfo.xml is intentionally lossy when a field has no equivalent in the target schema; the complete field mapping is maintained in `src/metadata/schema.ts`.
|
|
178
195
|
|
|
179
196
|
List fields that MetronInfo represents as an element with an optional `id` attribute (`genres`, `tags`, `characters`, `teams`, `locations`, `stories`, `reprints`) accept either a plain string or a `{ name, id? }` object (see [`MetronResource`](#types)); the `id` is preserved on round-trip through MetronInfo.xml and dropped when writing ComicInfo.xml, which has no equivalent concept.
|
|
180
197
|
|
|
@@ -266,7 +283,7 @@ const metadata = cah.xmlToMetadata(xml, 'ComicInfo');
|
|
|
266
283
|
|
|
267
284
|
### `validateMetadataXml(xml, schema)`
|
|
268
285
|
|
|
269
|
-
Validates an XML document against the bundled XSD for `'ComicInfo'` or `'MetronInfo'`, using [`xmllint-wasm`](https://www.npmjs.com/package/xmllint-wasm) (a WASM build of libxml2
|
|
286
|
+
Validates an XML document against the bundled XSD for `'ComicInfo'` or `'MetronInfo'`, using [`xmllint-wasm`](https://www.npmjs.com/package/xmllint-wasm) (a WASM build of libxml2; no native or Java dependency). Returns `{ valid, issues }`; `issues` is empty when the document conforms, otherwise it has one entry per schema violation with the offending element/attribute in `message` and, where available, a `line` number.
|
|
270
287
|
|
|
271
288
|
Because the validator implements XSD 1.0, it cannot check the two `<xs:assert>` business rules in MetronInfo.xml v1.1 (at most one primary `URL`, at most one primary `ID`); everything else in both schemas is enforced.
|
|
272
289
|
|
|
@@ -289,9 +306,21 @@ if (!valid) {
|
|
|
289
306
|
}
|
|
290
307
|
```
|
|
291
308
|
|
|
309
|
+
**ASAR is the one exception to everything below**: an asar archive never
|
|
310
|
+
stores metadata as a `ComicInfo.xml`/`MetronInfo.xml` entry. Instead,
|
|
311
|
+
`hasComicMetadata`/`readArchiveMetadata`/`addMetadataToArchive`/
|
|
312
|
+
`removeComicMetadata` all read/write a `comicMetadata` key embedded directly
|
|
313
|
+
in the asar container's own header (the small JSON block at the start of
|
|
314
|
+
every `.asar` file, alongside the header's existing `files` key), as plain
|
|
315
|
+
JSON; no XML involved. This is transparent to callers: the same four
|
|
316
|
+
functions work identically regardless of archive type, and `schema` still
|
|
317
|
+
selects `'ComicInfo'` vs `'MetronInfo'` field semantics either way. A
|
|
318
|
+
`ComicInfo.xml`/`MetronInfo.xml` _entry_ inside an asar archive is never
|
|
319
|
+
read as metadata by any of these functions.
|
|
320
|
+
|
|
292
321
|
### `hasComicMetadata(input)`
|
|
293
322
|
|
|
294
|
-
Checks for a root-level or nested `ComicInfo.xml` or `MetronInfo.xml` entry without parsing it. Returns `{ present, schema?, path? }`.
|
|
323
|
+
Checks for a root-level or nested `ComicInfo.xml` or `MetronInfo.xml` entry without parsing it (or, for asar, a `comicMetadata` header key). Returns `{ present, schema?, path?, bothPresent? }`. `path` is only set for a real archive entry (never for asar); `bothPresent` (asar only) is `true` when both `ComicInfo` and `MetronInfo` are present in the header.
|
|
295
324
|
|
|
296
325
|
**Options**
|
|
297
326
|
|
|
@@ -306,13 +335,14 @@ if (result.present) {
|
|
|
306
335
|
}
|
|
307
336
|
```
|
|
308
337
|
|
|
309
|
-
### `readArchiveMetadata(input)`
|
|
338
|
+
### `readArchiveMetadata(input, schema?)`
|
|
310
339
|
|
|
311
|
-
Finds and parses
|
|
340
|
+
Finds and parses metadata. With no `schema`, returns whichever is found first (asar checks ComicInfo before MetronInfo). Pass `schema` to read that one specifically regardless of preference. The only way to read a non-preferred schema back out of an asar archive that has both, since asar metadata isn't addressable by path. Returns `{ schema, metadata }`, or `null` when that metadata isn't present.
|
|
312
341
|
|
|
313
342
|
**Options**
|
|
314
343
|
|
|
315
344
|
- `input: ArchiveInput` - A filesystem path or archive `Buffer`.
|
|
345
|
+
- `schema?: 'ComicInfo' | 'MetronInfo'` - Read this specific schema instead of whichever is preferred.
|
|
316
346
|
|
|
317
347
|
**Example**
|
|
318
348
|
|
|
@@ -323,7 +353,7 @@ console.log(result?.metadata.title);
|
|
|
323
353
|
|
|
324
354
|
### `addMetadataToArchive(input, metadata, schema, options?)`
|
|
325
355
|
|
|
326
|
-
Adds a `ComicInfo.xml` or `MetronInfo.xml` entry to an archive. Existing metadata is rejected unless `overwrite: true` is supplied.
|
|
356
|
+
Adds a `ComicInfo.xml` or `MetronInfo.xml` entry to an archive (or, for asar, sets the `comicMetadata.{schema}` header key). Existing metadata is rejected unless `overwrite: true` is supplied.
|
|
327
357
|
|
|
328
358
|
**Options**
|
|
329
359
|
|
|
@@ -349,6 +379,23 @@ const withMetadata = await cah.addMetadataToArchive(
|
|
|
349
379
|
);
|
|
350
380
|
```
|
|
351
381
|
|
|
382
|
+
### `removeComicMetadata(input, schema, options?)`
|
|
383
|
+
|
|
384
|
+
Removes the embedded `schema` metadata (a `ComicInfo.xml`/`MetronInfo.xml` entry, or, for asar, the `comicMetadata.{schema}` header key). Throws `ArchiveFormatError` if that schema isn't present. Check `hasComicMetadata` first if you want a no-op instead.
|
|
385
|
+
|
|
386
|
+
**Options**
|
|
387
|
+
|
|
388
|
+
- `input: ArchiveInput` - A filesystem path or archive `Buffer`.
|
|
389
|
+
- `schema: 'ComicInfo' | 'MetronInfo'` - Metadata to remove.
|
|
390
|
+
- `options.tempDir?: string` - Temporary staging directory for ASAR or 7z operations.
|
|
391
|
+
- `options.output?: string | Writable` - Output destination. Without it, returns a `Buffer`.
|
|
392
|
+
|
|
393
|
+
**Example**
|
|
394
|
+
|
|
395
|
+
```js
|
|
396
|
+
const withoutMetronInfo = await cah.removeComicMetadata(comicBuffer, 'MetronInfo');
|
|
397
|
+
```
|
|
398
|
+
|
|
352
399
|
## Image conversion and detection
|
|
353
400
|
|
|
354
401
|
### `convertImageBuffer(image, format, options?)`
|
|
@@ -535,7 +582,7 @@ const digest = await cah.sha256Archive('/books/example.cbz');
|
|
|
535
582
|
|
|
536
583
|
### `benchmarkArchive(filePath, options?)`
|
|
537
584
|
|
|
538
|
-
Extracts a comic archive into the report's `source/` subdirectory, validates it contains image files, then generates one archive per writable container format (`zip`, `tar`, `asar`, `7z`) times benchmarked page image format (`webp`, `png`, `jpg` by default, or a subset via `options.imageFormats`)
|
|
585
|
+
Extracts a comic archive into the report's `source/` subdirectory, validates it contains image files, then generates one archive per writable container format (`zip`, `tar`, `asar`, `7z`) times benchmarked page image format (`webp`, `png`, `jpg` by default, or a subset via `options.imageFormats`). 12 variants in total by default, each containing only the original's XML and image entries, always re-encoded from the untouched source pages (never from a previously-converted variant). Benchmarks each variant's average archive-creation time (packaging only; images are re-encoded once per format before timing starts) and average random single-entry read time, writes the generated archives and a markdown report to a timestamped subdirectory, and returns the results.
|
|
539
586
|
|
|
540
587
|
Throws `ArchiveFormatError` (undetectable format) or `UnsupportedOperationError` (ACE) if `filePath` isn't extractable, `NoImagesFoundError` if the archive has no image files, `FilesystemAccessError` if `filePath` doesn't exist, and `RangeError` if `options.imageFormats` is empty or names an unsupported format.
|
|
541
588
|
|
|
@@ -605,6 +652,7 @@ The following types are exported for TypeScript consumers.
|
|
|
605
652
|
- `AddMetadataOptions` - Archive write options plus `overwrite?: boolean`.
|
|
606
653
|
- `RenameOptions` - Archive write options plus `start?: number` and `pad?: number`.
|
|
607
654
|
- `StripOptions` - Archive write options plus `extraKeepExtensions?: string[]`.
|
|
655
|
+
- `RemoveEntryOptions` - Archive write options (no additional fields).
|
|
608
656
|
- `ImageConvertOptions` - `webp?`, `jpeg?`, and `png?` format option groups.
|
|
609
657
|
- `WebpOptions` - `quality?`, `effort?`, and `smartSubsample?`.
|
|
610
658
|
- `JpegOptions` - `quality?`.
|
|
@@ -704,7 +752,7 @@ console.log(cah.joinResourceNames([{ name: 'Action', id: 'genre-1' }, 'Adventure
|
|
|
704
752
|
|
|
705
753
|
### Schema reference enums
|
|
706
754
|
|
|
707
|
-
Readonly value lists for the enumerated (`xs:enumeration`) types in each XSD. These are reference data, not enforced constraints
|
|
755
|
+
Readonly value lists for the enumerated (`xs:enumeration`) types in each XSD. These are reference data, not enforced constraints. Fields such as `ageRating` remain plain `string` so unrecognized or future values still round-trip.
|
|
708
756
|
|
|
709
757
|
- `COMIC_INFO_YES_NO_VALUES` - `'Unknown' | 'No' | 'Yes'` (ComicInfo `BlackAndWhite`).
|
|
710
758
|
- `COMIC_INFO_MANGA_VALUES` - `'Unknown' | 'No' | 'Yes' | 'YesAndRightToLeft'` (ComicInfo `Manga`).
|
|
@@ -713,7 +761,7 @@ Readonly value lists for the enumerated (`xs:enumeration`) types in each XSD. Th
|
|
|
713
761
|
- `METRON_FORMAT_VALUES` - MetronInfo `Series > Format` values, e.g. `'Single Issue'`, `'Trade Paperback'`.
|
|
714
762
|
- `METRON_INFORMATION_SOURCE_VALUES` - MetronInfo `IDS > ID` `source` attribute values, e.g. `'Comic Vine'`, `'Metron'`.
|
|
715
763
|
- `METRON_ROLE_VALUES` - MetronInfo `Credits > Credit > Roles > Role` values (a larger set than `COMIC_INFO_CREDIT_ROLES`).
|
|
716
|
-
- `METRON_AGE_RATING_VALUES` - MetronInfo `AgeRating` values
|
|
764
|
+
- `METRON_AGE_RATING_VALUES` - MetronInfo `AgeRating` values; a different set from ComicInfo's.
|
|
717
765
|
|
|
718
766
|
**Example**
|
|
719
767
|
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@clearmist-labs/comic-archive-handler",
|
|
3
|
-
"version": "1.
|
|
3
|
+
"version": "1.4.0",
|
|
4
4
|
"description": "Detect, convert, and manipulate comic book archives: container conversion, edit metadata, image re-encoding, perceptual + content hashing.",
|
|
5
5
|
"keywords": [
|
|
6
6
|
"7z",
|