fair-bioheaders 0.4.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,73 @@
1
+ CONTENTS
2
+
3
+ Public Domain Notice
4
+ Exceptions (for bundled 3rd-party code)
5
+ Copyright F.A.Q.
6
+
7
+
8
+ ==============================================================
9
+ PUBLIC DOMAIN NOTICE
10
+ United States Department of Agriculture
11
+ Agricultural Research Service
12
+
13
+ With the exception of certain third-party files summarized below, this
14
+ software is a "United States Government Work" under the terms of the
15
+ United States Copyright Act. It was written as part of the authors'
16
+ official duties as United States Government employees and thus cannot
17
+ be copyrighted. This software is freely available to the public for
18
+ use. The United States Department of Agriculture, Agricultural
19
+ Research Service (USDA - ARS) and the U.S. Government have not placed
20
+ any restriction on its use or reproduction.
21
+
22
+ Although all reasonable efforts have been taken to ensure the accuracy
23
+ and reliability of the software and data, the USDA ARS and the U.S.
24
+ Government do not and cannot warrant the performance or results tha may
25
+ be obtained by using this software or data. The USDA ARS and the U.S.
26
+ Government disclaim all warranties, express or implied, including
27
+ warranties of performance, merchantability or fitness for any particular
28
+ purpose.
29
+
30
+ Please cite the authors in any work or product based on this material.
31
+
32
+
33
+ ==============================================================
34
+ EXCEPTIONS (in all cases excluding USDA-ARS-written makefiles):
35
+
36
+ Location:
37
+ Author:
38
+ License:
39
+
40
+
41
+ ==============================================================
42
+ Copyright F.A.Q.
43
+
44
+
45
+ --------------------------------------------------------------
46
+ Q. Our product makes use of the USDA - ARS source code, and we made changes
47
+ and additions to that version of the USDA - ARS code to better fit it to
48
+ our needs. Can we copyright the code, and how?
49
+
50
+ A. You can copyright only the *changes* or the *additions* you made to the
51
+ NCBI source code. You should identify unambiguously those sections of
52
+ the code that were modified, e.g. by commenting any changes you made
53
+ in the code you distribute. Therefore, your license has to make clear
54
+ to users that your product is a combination of code that is public domain
55
+ within the U.S. (but may be subject to copyright by the U.S. in foreign
56
+ countries) and code that has been created or modified by you.
57
+
58
+ --------------------------------------------------------------
59
+ Q. Can we (re)license all or part of the USDA - ARS source code?
60
+
61
+ A. No, you cannot license or relicense the source code written by USDA - ARS
62
+ since you cannot claim any copyright in the software that was developed
63
+ at USDA - ARS as a 'government work' and consequently is in the public
64
+ domain within the U.S.
65
+
66
+ --------------------------------------------------------------
67
+ Q. What if these copyright guidelines are not clear enough or are not
68
+ applicable to my particular case?
69
+
70
+ A. Contact us. Send your questions to 'answers@usda.gov'.
71
+ --------------------------------------------------------------
72
+
73
+ This file was modified from the NCBI Boilerplate LICENSE file
@@ -0,0 +1,261 @@
1
+ Metadata-Version: 2.4
2
+ Name: fair-bioheaders
3
+ Version: 0.4.0
4
+ Summary: Convert and validate FAIR-bioHeaders (FHR) metadata in JSON, YAML, FASTA, GFA, and HTML (formerly the fhr package)
5
+ License: USDA-ARS
6
+ License-File: LICENSE
7
+ Keywords: FAIR,FHR,genome,metadata,fasta,gfa
8
+ Author: David Molik
9
+ Author-email: david.molik@usda.gov
10
+ Requires-Python: >=3.9,<4.0
11
+ Classifier: License :: Other/Proprietary License
12
+ Classifier: Programming Language :: Python :: 3
13
+ Classifier: Programming Language :: Python :: 3.9
14
+ Classifier: Programming Language :: Python :: 3.10
15
+ Classifier: Programming Language :: Python :: 3.11
16
+ Classifier: Programming Language :: Python :: 3.12
17
+ Classifier: Programming Language :: Python :: 3.13
18
+ Classifier: Programming Language :: Python :: 3.14
19
+ Requires-Dist: jsonschema[format] (>=4.21.1,<5.0.0)
20
+ Requires-Dist: pyyaml (>=6.0.1,<7.0.0)
21
+ Project-URL: Bug Tracker, https://github.com/FAIR-bioHeaders/FAIR-bioHeaders-Tools/issues
22
+ Project-URL: Changelog, https://github.com/FAIR-bioHeaders/FAIR-bioHeaders-Tools/blob/main/CHANGELOG.md
23
+ Project-URL: Documentation, https://github.com/FAIR-bioHeaders/FHR-Specification
24
+ Project-URL: Homepage, https://github.com/FAIR-bioHeaders/FAIR-bioHeaders-Tools
25
+ Project-URL: Repository, https://github.com/FAIR-bioHeaders/FAIR-bioHeaders-Tools
26
+ Description-Content-Type: text/markdown
27
+
28
+ # FAIR-bioHeaders Tools
29
+
30
+ [![Tests](https://github.com/FAIR-bioHeaders/FAIR-bioHeaders-Tools/actions/workflows/pytest.yaml/badge.svg?branch=main)](https://github.com/FAIR-bioHeaders/FAIR-bioHeaders-Tools/actions/workflows/pytest.yaml)
31
+ [![Specification DOI](https://img.shields.io/badge/Specification_DOI-10.5281%2Fzenodo.6762549-blue)](https://doi.org/10.5281/zenodo.6762549)
32
+ [![File Converter DOI](https://img.shields.io/badge/File_Converter_DOI-10.5281%2Fzenodo.6762547-blue)](https://doi.org/10.5281/zenodo.6762547)
33
+
34
+ Convert and validate FAIR-bioHeaders Reference genome (FHR) metadata in JSON,
35
+ YAML, FASTA, GFA, and HTML microdata. See
36
+ [FHR-Specification](https://github.com/FAIR-bioHeaders/FHR-Specification) for the
37
+ schema and metadata design. This is the `fair-bioheaders` package, version
38
+ **0.4.0**, formerly the FHR File Converter (`fhr`); see
39
+ [Renamed from fhr](#renamed-from-fhr).
40
+
41
+ ## Install
42
+
43
+ For a published release:
44
+
45
+ ```bash
46
+ python -m pip install fair-bioheaders
47
+ ```
48
+
49
+ For this checkout and its development checks (Python 3.9 or later):
50
+
51
+ ```bash
52
+ python -m pip install poetry
53
+ poetry install
54
+ poetry run pytest
55
+ poetry run ruff check .
56
+ poetry run isort . --check-only
57
+ poetry run black . --check
58
+ ```
59
+
60
+ ## Commands
61
+
62
+ ```bash
63
+ bioheaders convert examples/example.fhr.yaml /tmp/example.fhr.json
64
+ bioheaders validate /tmp/example.fhr.json
65
+ bioheaders combine examples/example.fhr.yaml genome.fasta -o genome.fhr.fasta
66
+ bioheaders verify genome.fhr.fasta
67
+ bioheaders strip genome.fhr.fasta genome.stripped.fasta
68
+ bioheaders combine examples/example.fhr.yaml assembly.gfa -o assembly.fhr.gfa
69
+ bioheaders verify assembly.fhr.gfa
70
+ bioheaders strip assembly.fhr.gfa assembly.stripped.gfa
71
+ ```
72
+
73
+ `combine`, `strip`, and `verify` (also called `checksum`) take the FASTA or GFA
74
+ type from the sequence file's extension, ignoring `.gz` or `.bgz`; give
75
+ `--type fasta` or `--type gfa` for stdin or other extensions. `bioheaders
76
+ --version` prints the version and `bioheaders COMMAND --help` describes each
77
+ command. The original commands remain available with the same options and output:
78
+
79
+ | `bioheaders` command | Original command |
80
+ | --- | --- |
81
+ | `convert INPUT OUTPUT` | `fhr-convert` |
82
+ | `validate INPUT` | `fhr-validate` |
83
+ | `combine METADATA SEQUENCE` | `fhr-fasta-combine`, `fhr-gfa-combine` |
84
+ | `strip INPUT [OUTPUT]` | `fhr-fasta-strip`, `fhr-gfa-strip` |
85
+ | `verify INPUT` | `fhr-fasta-validate`, `fhr-gfa-validate` |
86
+
87
+ `convert INPUT OUTPUT` detects `.json`, `.yaml`/`.yml`, `.fasta`/`.fa`/`.fna`,
88
+ `.gfa`, and `.html` from extensions. It validates metadata before writing output;
89
+ FASTA/GFA output contains a metadata header only. To include sequence data, use
90
+ `combine METADATA SEQUENCE [-o OUTPUT]`, whose default output is `SEQUENCE.fhr.fasta`
91
+ or `SEQUENCE.fhr.gfa` with the last extension replaced. Existing FHR header lines
92
+ are replaced. Inputs are not overwritten by sequence helpers.
93
+
94
+ `strip INPUT [OUTPUT]` writes to stdout if OUTPUT is omitted. Only FHR-prefixed
95
+ lines are removed; other bytes, including ordinary comments and CRLF endings,
96
+ are preserved. `validate` checks metadata, while `verify` (`fhr-fasta-validate`,
97
+ `fhr-gfa-validate`) also verifies the exact-byte file checksum. Failures exit
98
+ with status 1.
99
+
100
+ FASTA/GFA commands stream their input in 1 MiB chunks, so memory use does not grow
101
+ with file size (about 35 MB peak for a 1 GB FASTA). Validate reads the file once;
102
+ combine reads it twice. FHR header lines are limited to 16 MiB in total. Outputs
103
+ are written to a temporary file in the destination directory and then renamed,
104
+ so a failed command leaves no partial output.
105
+
106
+ ### Compressed files and pipes
107
+
108
+ ```bash
109
+ bioheaders combine examples/example.fhr.yaml genome.fa.gz # writes genome.fhr.fasta.gz
110
+ bioheaders verify genome.fhr.fasta.gz
111
+ bioheaders strip genome.fhr.fasta.gz genome.stripped.fa.gz
112
+ zcat genome.fa.gz | bioheaders combine --type fasta examples/example.fhr.yaml - > genome.fhr.fasta
113
+ curl -sL https://example.org/genome.fhr.fasta.gz | bioheaders verify --type fasta -
114
+ bioheaders convert genome.fhr.fasta.gz - --to json
115
+ bioheaders convert - metadata.yaml --from json < metadata.json
116
+ ```
117
+
118
+ Gzip input, including multi-member gzip and BGZF, is recognized by its magic
119
+ bytes rather than its extension and decompressed as it streams. The checksum
120
+ covers the decompressed FASTA/GFA bytes, so `genome.fa` and any gzip or BGZF
121
+ compression of it have the same checksum. A corrupt or truncated gzip file is an
122
+ error. Formats are taken from the extension before `.gz` or `.bgz`.
123
+
124
+ Outputs whose path ends in `.gz` or `.bgz` are written as BGZF, which `gzip`,
125
+ `zcat`, `bgzip`, and htslib read. A combine without `-o` keeps the input's
126
+ compression extension. Other outputs, and stdout, are never compressed. Note that
127
+ `samtools faidx` rejects FASTA files containing `;` lines, compressed or not, so
128
+ index a stripped copy for random access.
129
+
130
+ `-` reads stdin or writes stdout: the input of verify, strip, combine (the
131
+ sequence), convert, and validate, and the output of strip, convert, and
132
+ `combine -o`. Combine of stdin writes stdout by default. Convert and validate
133
+ need `--from FORMAT` for stdin and convert needs `--to FORMAT` for stdout
134
+ (`json`, `yaml`, `fasta`, `gfa`, or `html`); the `bioheaders` sequence commands
135
+ need `--type` for stdin. Stdin is read
136
+ once: combine spools the stripped sequence to a temporary file in `TMPDIR` (as
137
+ large as the sequence) while hashing it. Strip from stdin to stdout writes as it
138
+ reads, so an error late in the input can follow partial output; check the exit
139
+ status. Strip of a file to stdout still checks the whole file before writing.
140
+
141
+ ## Python
142
+
143
+ ```python
144
+ from bioheaders import fhr
145
+
146
+ metadata = fhr()
147
+ with open("examples/example.fhr.yaml", encoding="utf-8") as stream:
148
+ metadata.input_yaml(stream)
149
+ metadata.fhr_validate()
150
+ print(metadata.output_json())
151
+ ```
152
+
153
+ Input methods accept text, UTF-8 bytes, or readable streams. Optional fields stay
154
+ absent until supplied. The mapping is preserved across supported format round
155
+ trips; HTML values are escaped and explicitly typed. Keywords can initialize an
156
+ instance (`fhr(genome="example")`), and fields remain accessible as attributes.
157
+ The installed schema is loaded from package resources rather than the working
158
+ directory, so commands work outside the checkout. No schema is downloaded at runtime.
159
+ `bioheaders.cli` holds the command implementations and the byte-level helpers
160
+ (`checksum`, `combine`, `strip_header`, `open_input`, `write_to`, and others).
161
+
162
+ ## Renamed from fhr
163
+
164
+ In 0.4.0 the PyPI package `fhr` was renamed `fair-bioheaders`, the Python package
165
+ `fhr` was renamed `bioheaders`, and the repository moved from FHR-File-Converter
166
+ to [FAIR-bioHeaders-Tools](https://github.com/FAIR-bioHeaders/FAIR-bioHeaders-Tools)
167
+ (old links redirect). The new names leave room for other FAIR-bioHeaders header
168
+ types. Nothing that worked with 0.3 stops working:
169
+
170
+ - The eight `fhr-*` commands are unchanged and print no deprecation notices, so
171
+ scripts and pipelines that read their output or stderr keep working. They are
172
+ not deprecated, but new scripts can use `bioheaders`.
173
+ - `import fhr`, `from fhr import fhr`, and `from fhr.cli import ...` still work:
174
+ `fhr` and `fhr.cli` are the same module objects as `bioheaders` and
175
+ `bioheaders.cli`. Importing `fhr` emits one `DeprecationWarning`; replace
176
+ `fhr` with `bioheaders` in imports. The `fhr` metadata class keeps its name.
177
+ - `pip install fhr` installs `fhr` 0.4.0, a compatibility package without code
178
+ that requires the same version of `fair-bioheaders`, so existing requirements
179
+ and `pip install -U fhr` keep working. Prefer `fair-bioheaders` in new
180
+ requirements. The compatibility package also installs the `fhr-*` commands, so
181
+ if you uninstall `fhr` 0.4 later, restore them with
182
+ `python -m pip install --force-reinstall --no-deps fair-bioheaders`.
183
+ - In an environment that still has `fhr` 0.3, run `python -m pip uninstall fhr`
184
+ before `python -m pip install fair-bioheaders` (or upgrade with
185
+ `python -m pip install -U fhr`); otherwise the old `fhr` package files shadow
186
+ the new compatibility module.
187
+ - Citations and Zenodo DOIs are unchanged.
188
+
189
+ Checksum helpers require SHA-512/256 support in the Python build. Some Apple
190
+ system Python builds omit it; use an OpenSSL-enabled Python distribution.
191
+ Metadata conversion and validation do not require that hash implementation.
192
+
193
+ ## v0.3 compatibility and identity
194
+
195
+ - Required fields and `schemaVersion: 1` remain unchanged since v0.3.0.
196
+ - `assemblySoftware` accepts a legacy string or optional structured software
197
+ objects with name, URI, version, and command options. `assemblyProtocol` is a URI.
198
+ - `vitalStats.N90` is base pairs; `vitalStats.gcContent` is 0–100 percent.
199
+ - Optional `seqcol_id` stores a supplied 32-character base64url top-level refget
200
+ digest. The converter preserves it; it does not compute or verify SeqCol identity.
201
+ - Checksums use base64 SHA-512/256 of all original file bytes except the one scalar
202
+ checksum header line, including its newline. Metadata is covered. MD5 hex and
203
+ payload-only checksums from old examples are not valid v0.3 checksums; recombine
204
+ metadata with the original sequence file to calculate the new value.
205
+ - The checksum value must be a single-line scalar on the checksum line. Header
206
+ lines must be UTF-8 and must not contain U+0085, U+2028, or U+2029. FASTA/GFA
207
+ files must not begin with a UTF-8 byte order mark. Other sequence bytes may use
208
+ any encoding.
209
+ - Metadata must be JSON-compatible: duplicate keys, YAML anchors, aliases, and
210
+ merge keys are rejected.
211
+ - `;~`/`#~` lines must form the leading header block, before the first FASTA `>`
212
+ line or the first GFA record line; ordinary comments and blank lines may be
213
+ mixed in. A `;~`/`#~` line after sequence data, including a concatenated
214
+ second file, is an error.
215
+ - HTML exports use FHR item scopes and `data-fhr-type` annotations for lossless
216
+ arrays, numbers, objects, and strings. External microdata must represent nested
217
+ items properly; incomplete legacy markup may need regeneration.
218
+ - Incomplete constructors now emit missing fields rather than empty defaults.
219
+ Validate after loading to obtain actionable schema errors. Obsolete positional
220
+ constructor arguments should be converted to keywords.
221
+
222
+ See the [format reference](https://github.com/FAIR-bioHeaders/FHR-Specification/blob/main/docs/FORMAT.md)
223
+ and [release notes](CHANGELOG.md). JSON/YAML/HTML example identifiers are synthetic.
224
+ The FASTA/GFA fixtures contain verified FHR checksums; their SeqCol IDs are placeholders.
225
+
226
+ ## Docker
227
+
228
+ ```bash
229
+ docker build -t fair-bioheaders-tools .
230
+ docker run --rm fair-bioheaders-tools --help
231
+ ```
232
+
233
+ The image runs `fhr-convert`; use `--entrypoint bioheaders` for the other
234
+ commands. Mount inputs/outputs into the container for conversion. See
235
+ [CONTRIBUTING](CONTRIBUTING.md), [CODE_OF_CONDUCT](CODE_OF_CONDUCT.md),
236
+ [SECURITY](SECURITY.md), and [AGENTS](AGENTS.md) for project guidance.
237
+
238
+ ## Citing FHR
239
+
240
+ Chicago bibliography entries are used below. Cite the published paper for a
241
+ general description of FHR; cite the specification or converter when using that
242
+ resource directly. The software and specification links are concept DOIs; for a
243
+ specific release, use the corresponding version DOI from Zenodo. Authors and
244
+ years follow the records resolved by the concept DOIs at the v0.3 documentation
245
+ update, and can change as later records are published.
246
+
247
+ ### Published paper
248
+
249
+ Wright, Adam, Mark D. Wilkinson, Christopher Mungall, Scott Cain, Stephen Richards, Paul Sternberg, Ellen Provin, Jonathan L. Jacobs, Scott Geib, Daniela Raciti, Karen Yook, Lincoln Stein, and David C. Molik. “FAIR Header Reference Genome: A TRUSTworthy Standard.” *Briefings in Bioinformatics* 25, no. 3 (2024): bbae122. https://doi.org/10.1093/bib/bbae122.
250
+
251
+ ### Specification
252
+
253
+ Molik, David. *FHR Specification*. Data set. 2022. https://doi.org/10.5281/zenodo.6762549.
254
+
255
+ ### Converter
256
+
257
+ Molik, David, and Adam Wright. *FHR File Converter*. Computer software. 2024. https://doi.org/10.5281/zenodo.6762547.
258
+
259
+ Machine-readable entries are maintained in
260
+ [FHR-Citation](https://github.com/FAIR-bioHeaders/FHR-Citation/blob/main/citation.bib).
261
+
@@ -0,0 +1,233 @@
1
+ # FAIR-bioHeaders Tools
2
+
3
+ [![Tests](https://github.com/FAIR-bioHeaders/FAIR-bioHeaders-Tools/actions/workflows/pytest.yaml/badge.svg?branch=main)](https://github.com/FAIR-bioHeaders/FAIR-bioHeaders-Tools/actions/workflows/pytest.yaml)
4
+ [![Specification DOI](https://img.shields.io/badge/Specification_DOI-10.5281%2Fzenodo.6762549-blue)](https://doi.org/10.5281/zenodo.6762549)
5
+ [![File Converter DOI](https://img.shields.io/badge/File_Converter_DOI-10.5281%2Fzenodo.6762547-blue)](https://doi.org/10.5281/zenodo.6762547)
6
+
7
+ Convert and validate FAIR-bioHeaders Reference genome (FHR) metadata in JSON,
8
+ YAML, FASTA, GFA, and HTML microdata. See
9
+ [FHR-Specification](https://github.com/FAIR-bioHeaders/FHR-Specification) for the
10
+ schema and metadata design. This is the `fair-bioheaders` package, version
11
+ **0.4.0**, formerly the FHR File Converter (`fhr`); see
12
+ [Renamed from fhr](#renamed-from-fhr).
13
+
14
+ ## Install
15
+
16
+ For a published release:
17
+
18
+ ```bash
19
+ python -m pip install fair-bioheaders
20
+ ```
21
+
22
+ For this checkout and its development checks (Python 3.9 or later):
23
+
24
+ ```bash
25
+ python -m pip install poetry
26
+ poetry install
27
+ poetry run pytest
28
+ poetry run ruff check .
29
+ poetry run isort . --check-only
30
+ poetry run black . --check
31
+ ```
32
+
33
+ ## Commands
34
+
35
+ ```bash
36
+ bioheaders convert examples/example.fhr.yaml /tmp/example.fhr.json
37
+ bioheaders validate /tmp/example.fhr.json
38
+ bioheaders combine examples/example.fhr.yaml genome.fasta -o genome.fhr.fasta
39
+ bioheaders verify genome.fhr.fasta
40
+ bioheaders strip genome.fhr.fasta genome.stripped.fasta
41
+ bioheaders combine examples/example.fhr.yaml assembly.gfa -o assembly.fhr.gfa
42
+ bioheaders verify assembly.fhr.gfa
43
+ bioheaders strip assembly.fhr.gfa assembly.stripped.gfa
44
+ ```
45
+
46
+ `combine`, `strip`, and `verify` (also called `checksum`) take the FASTA or GFA
47
+ type from the sequence file's extension, ignoring `.gz` or `.bgz`; give
48
+ `--type fasta` or `--type gfa` for stdin or other extensions. `bioheaders
49
+ --version` prints the version and `bioheaders COMMAND --help` describes each
50
+ command. The original commands remain available with the same options and output:
51
+
52
+ | `bioheaders` command | Original command |
53
+ | --- | --- |
54
+ | `convert INPUT OUTPUT` | `fhr-convert` |
55
+ | `validate INPUT` | `fhr-validate` |
56
+ | `combine METADATA SEQUENCE` | `fhr-fasta-combine`, `fhr-gfa-combine` |
57
+ | `strip INPUT [OUTPUT]` | `fhr-fasta-strip`, `fhr-gfa-strip` |
58
+ | `verify INPUT` | `fhr-fasta-validate`, `fhr-gfa-validate` |
59
+
60
+ `convert INPUT OUTPUT` detects `.json`, `.yaml`/`.yml`, `.fasta`/`.fa`/`.fna`,
61
+ `.gfa`, and `.html` from extensions. It validates metadata before writing output;
62
+ FASTA/GFA output contains a metadata header only. To include sequence data, use
63
+ `combine METADATA SEQUENCE [-o OUTPUT]`, whose default output is `SEQUENCE.fhr.fasta`
64
+ or `SEQUENCE.fhr.gfa` with the last extension replaced. Existing FHR header lines
65
+ are replaced. Inputs are not overwritten by sequence helpers.
66
+
67
+ `strip INPUT [OUTPUT]` writes to stdout if OUTPUT is omitted. Only FHR-prefixed
68
+ lines are removed; other bytes, including ordinary comments and CRLF endings,
69
+ are preserved. `validate` checks metadata, while `verify` (`fhr-fasta-validate`,
70
+ `fhr-gfa-validate`) also verifies the exact-byte file checksum. Failures exit
71
+ with status 1.
72
+
73
+ FASTA/GFA commands stream their input in 1 MiB chunks, so memory use does not grow
74
+ with file size (about 35 MB peak for a 1 GB FASTA). Validate reads the file once;
75
+ combine reads it twice. FHR header lines are limited to 16 MiB in total. Outputs
76
+ are written to a temporary file in the destination directory and then renamed,
77
+ so a failed command leaves no partial output.
78
+
79
+ ### Compressed files and pipes
80
+
81
+ ```bash
82
+ bioheaders combine examples/example.fhr.yaml genome.fa.gz # writes genome.fhr.fasta.gz
83
+ bioheaders verify genome.fhr.fasta.gz
84
+ bioheaders strip genome.fhr.fasta.gz genome.stripped.fa.gz
85
+ zcat genome.fa.gz | bioheaders combine --type fasta examples/example.fhr.yaml - > genome.fhr.fasta
86
+ curl -sL https://example.org/genome.fhr.fasta.gz | bioheaders verify --type fasta -
87
+ bioheaders convert genome.fhr.fasta.gz - --to json
88
+ bioheaders convert - metadata.yaml --from json < metadata.json
89
+ ```
90
+
91
+ Gzip input, including multi-member gzip and BGZF, is recognized by its magic
92
+ bytes rather than its extension and decompressed as it streams. The checksum
93
+ covers the decompressed FASTA/GFA bytes, so `genome.fa` and any gzip or BGZF
94
+ compression of it have the same checksum. A corrupt or truncated gzip file is an
95
+ error. Formats are taken from the extension before `.gz` or `.bgz`.
96
+
97
+ Outputs whose path ends in `.gz` or `.bgz` are written as BGZF, which `gzip`,
98
+ `zcat`, `bgzip`, and htslib read. A combine without `-o` keeps the input's
99
+ compression extension. Other outputs, and stdout, are never compressed. Note that
100
+ `samtools faidx` rejects FASTA files containing `;` lines, compressed or not, so
101
+ index a stripped copy for random access.
102
+
103
+ `-` reads stdin or writes stdout: the input of verify, strip, combine (the
104
+ sequence), convert, and validate, and the output of strip, convert, and
105
+ `combine -o`. Combine of stdin writes stdout by default. Convert and validate
106
+ need `--from FORMAT` for stdin and convert needs `--to FORMAT` for stdout
107
+ (`json`, `yaml`, `fasta`, `gfa`, or `html`); the `bioheaders` sequence commands
108
+ need `--type` for stdin. Stdin is read
109
+ once: combine spools the stripped sequence to a temporary file in `TMPDIR` (as
110
+ large as the sequence) while hashing it. Strip from stdin to stdout writes as it
111
+ reads, so an error late in the input can follow partial output; check the exit
112
+ status. Strip of a file to stdout still checks the whole file before writing.
113
+
114
+ ## Python
115
+
116
+ ```python
117
+ from bioheaders import fhr
118
+
119
+ metadata = fhr()
120
+ with open("examples/example.fhr.yaml", encoding="utf-8") as stream:
121
+ metadata.input_yaml(stream)
122
+ metadata.fhr_validate()
123
+ print(metadata.output_json())
124
+ ```
125
+
126
+ Input methods accept text, UTF-8 bytes, or readable streams. Optional fields stay
127
+ absent until supplied. The mapping is preserved across supported format round
128
+ trips; HTML values are escaped and explicitly typed. Keywords can initialize an
129
+ instance (`fhr(genome="example")`), and fields remain accessible as attributes.
130
+ The installed schema is loaded from package resources rather than the working
131
+ directory, so commands work outside the checkout. No schema is downloaded at runtime.
132
+ `bioheaders.cli` holds the command implementations and the byte-level helpers
133
+ (`checksum`, `combine`, `strip_header`, `open_input`, `write_to`, and others).
134
+
135
+ ## Renamed from fhr
136
+
137
+ In 0.4.0 the PyPI package `fhr` was renamed `fair-bioheaders`, the Python package
138
+ `fhr` was renamed `bioheaders`, and the repository moved from FHR-File-Converter
139
+ to [FAIR-bioHeaders-Tools](https://github.com/FAIR-bioHeaders/FAIR-bioHeaders-Tools)
140
+ (old links redirect). The new names leave room for other FAIR-bioHeaders header
141
+ types. Nothing that worked with 0.3 stops working:
142
+
143
+ - The eight `fhr-*` commands are unchanged and print no deprecation notices, so
144
+ scripts and pipelines that read their output or stderr keep working. They are
145
+ not deprecated, but new scripts can use `bioheaders`.
146
+ - `import fhr`, `from fhr import fhr`, and `from fhr.cli import ...` still work:
147
+ `fhr` and `fhr.cli` are the same module objects as `bioheaders` and
148
+ `bioheaders.cli`. Importing `fhr` emits one `DeprecationWarning`; replace
149
+ `fhr` with `bioheaders` in imports. The `fhr` metadata class keeps its name.
150
+ - `pip install fhr` installs `fhr` 0.4.0, a compatibility package without code
151
+ that requires the same version of `fair-bioheaders`, so existing requirements
152
+ and `pip install -U fhr` keep working. Prefer `fair-bioheaders` in new
153
+ requirements. The compatibility package also installs the `fhr-*` commands, so
154
+ if you uninstall `fhr` 0.4 later, restore them with
155
+ `python -m pip install --force-reinstall --no-deps fair-bioheaders`.
156
+ - In an environment that still has `fhr` 0.3, run `python -m pip uninstall fhr`
157
+ before `python -m pip install fair-bioheaders` (or upgrade with
158
+ `python -m pip install -U fhr`); otherwise the old `fhr` package files shadow
159
+ the new compatibility module.
160
+ - Citations and Zenodo DOIs are unchanged.
161
+
162
+ Checksum helpers require SHA-512/256 support in the Python build. Some Apple
163
+ system Python builds omit it; use an OpenSSL-enabled Python distribution.
164
+ Metadata conversion and validation do not require that hash implementation.
165
+
166
+ ## v0.3 compatibility and identity
167
+
168
+ - Required fields and `schemaVersion: 1` remain unchanged since v0.3.0.
169
+ - `assemblySoftware` accepts a legacy string or optional structured software
170
+ objects with name, URI, version, and command options. `assemblyProtocol` is a URI.
171
+ - `vitalStats.N90` is base pairs; `vitalStats.gcContent` is 0–100 percent.
172
+ - Optional `seqcol_id` stores a supplied 32-character base64url top-level refget
173
+ digest. The converter preserves it; it does not compute or verify SeqCol identity.
174
+ - Checksums use base64 SHA-512/256 of all original file bytes except the one scalar
175
+ checksum header line, including its newline. Metadata is covered. MD5 hex and
176
+ payload-only checksums from old examples are not valid v0.3 checksums; recombine
177
+ metadata with the original sequence file to calculate the new value.
178
+ - The checksum value must be a single-line scalar on the checksum line. Header
179
+ lines must be UTF-8 and must not contain U+0085, U+2028, or U+2029. FASTA/GFA
180
+ files must not begin with a UTF-8 byte order mark. Other sequence bytes may use
181
+ any encoding.
182
+ - Metadata must be JSON-compatible: duplicate keys, YAML anchors, aliases, and
183
+ merge keys are rejected.
184
+ - `;~`/`#~` lines must form the leading header block, before the first FASTA `>`
185
+ line or the first GFA record line; ordinary comments and blank lines may be
186
+ mixed in. A `;~`/`#~` line after sequence data, including a concatenated
187
+ second file, is an error.
188
+ - HTML exports use FHR item scopes and `data-fhr-type` annotations for lossless
189
+ arrays, numbers, objects, and strings. External microdata must represent nested
190
+ items properly; incomplete legacy markup may need regeneration.
191
+ - Incomplete constructors now emit missing fields rather than empty defaults.
192
+ Validate after loading to obtain actionable schema errors. Obsolete positional
193
+ constructor arguments should be converted to keywords.
194
+
195
+ See the [format reference](https://github.com/FAIR-bioHeaders/FHR-Specification/blob/main/docs/FORMAT.md)
196
+ and [release notes](CHANGELOG.md). JSON/YAML/HTML example identifiers are synthetic.
197
+ The FASTA/GFA fixtures contain verified FHR checksums; their SeqCol IDs are placeholders.
198
+
199
+ ## Docker
200
+
201
+ ```bash
202
+ docker build -t fair-bioheaders-tools .
203
+ docker run --rm fair-bioheaders-tools --help
204
+ ```
205
+
206
+ The image runs `fhr-convert`; use `--entrypoint bioheaders` for the other
207
+ commands. Mount inputs/outputs into the container for conversion. See
208
+ [CONTRIBUTING](CONTRIBUTING.md), [CODE_OF_CONDUCT](CODE_OF_CONDUCT.md),
209
+ [SECURITY](SECURITY.md), and [AGENTS](AGENTS.md) for project guidance.
210
+
211
+ ## Citing FHR
212
+
213
+ Chicago bibliography entries are used below. Cite the published paper for a
214
+ general description of FHR; cite the specification or converter when using that
215
+ resource directly. The software and specification links are concept DOIs; for a
216
+ specific release, use the corresponding version DOI from Zenodo. Authors and
217
+ years follow the records resolved by the concept DOIs at the v0.3 documentation
218
+ update, and can change as later records are published.
219
+
220
+ ### Published paper
221
+
222
+ Wright, Adam, Mark D. Wilkinson, Christopher Mungall, Scott Cain, Stephen Richards, Paul Sternberg, Ellen Provin, Jonathan L. Jacobs, Scott Geib, Daniela Raciti, Karen Yook, Lincoln Stein, and David C. Molik. “FAIR Header Reference Genome: A TRUSTworthy Standard.” *Briefings in Bioinformatics* 25, no. 3 (2024): bbae122. https://doi.org/10.1093/bib/bbae122.
223
+
224
+ ### Specification
225
+
226
+ Molik, David. *FHR Specification*. Data set. 2022. https://doi.org/10.5281/zenodo.6762549.
227
+
228
+ ### Converter
229
+
230
+ Molik, David, and Adam Wright. *FHR File Converter*. Computer software. 2024. https://doi.org/10.5281/zenodo.6762547.
231
+
232
+ Machine-readable entries are maintained in
233
+ [FHR-Citation](https://github.com/FAIR-bioHeaders/FHR-Citation/blob/main/citation.bib).