fair-bioheaders 0.4.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- fair_bioheaders-0.4.0/LICENSE +73 -0
- fair_bioheaders-0.4.0/PKG-INFO +261 -0
- fair_bioheaders-0.4.0/README.md +233 -0
- fair_bioheaders-0.4.0/bioheaders/__init__.py +678 -0
- fair_bioheaders-0.4.0/bioheaders/__main__.py +6 -0
- fair_bioheaders-0.4.0/bioheaders/cli.py +828 -0
- fair_bioheaders-0.4.0/bioheaders/fhr_schema.json +266 -0
- fair_bioheaders-0.4.0/fhr.py +30 -0
- fair_bioheaders-0.4.0/pyproject.toml +71 -0
|
@@ -0,0 +1,73 @@
|
|
|
1
|
+
CONTENTS
|
|
2
|
+
|
|
3
|
+
Public Domain Notice
|
|
4
|
+
Exceptions (for bundled 3rd-party code)
|
|
5
|
+
Copyright F.A.Q.
|
|
6
|
+
|
|
7
|
+
|
|
8
|
+
==============================================================
|
|
9
|
+
PUBLIC DOMAIN NOTICE
|
|
10
|
+
United States Department of Agriculture
|
|
11
|
+
Agricultural Research Service
|
|
12
|
+
|
|
13
|
+
With the exception of certain third-party files summarized below, this
|
|
14
|
+
software is a "United States Government Work" under the terms of the
|
|
15
|
+
United States Copyright Act. It was written as part of the authors'
|
|
16
|
+
official duties as United States Government employees and thus cannot
|
|
17
|
+
be copyrighted. This software is freely available to the public for
|
|
18
|
+
use. The United States Department of Agriculture, Agricultural
|
|
19
|
+
Research Service (USDA - ARS) and the U.S. Government have not placed
|
|
20
|
+
any restriction on its use or reproduction.
|
|
21
|
+
|
|
22
|
+
Although all reasonable efforts have been taken to ensure the accuracy
|
|
23
|
+
and reliability of the software and data, the USDA ARS and the U.S.
|
|
24
|
+
Government do not and cannot warrant the performance or results tha may
|
|
25
|
+
be obtained by using this software or data. The USDA ARS and the U.S.
|
|
26
|
+
Government disclaim all warranties, express or implied, including
|
|
27
|
+
warranties of performance, merchantability or fitness for any particular
|
|
28
|
+
purpose.
|
|
29
|
+
|
|
30
|
+
Please cite the authors in any work or product based on this material.
|
|
31
|
+
|
|
32
|
+
|
|
33
|
+
==============================================================
|
|
34
|
+
EXCEPTIONS (in all cases excluding USDA-ARS-written makefiles):
|
|
35
|
+
|
|
36
|
+
Location:
|
|
37
|
+
Author:
|
|
38
|
+
License:
|
|
39
|
+
|
|
40
|
+
|
|
41
|
+
==============================================================
|
|
42
|
+
Copyright F.A.Q.
|
|
43
|
+
|
|
44
|
+
|
|
45
|
+
--------------------------------------------------------------
|
|
46
|
+
Q. Our product makes use of the USDA - ARS source code, and we made changes
|
|
47
|
+
and additions to that version of the USDA - ARS code to better fit it to
|
|
48
|
+
our needs. Can we copyright the code, and how?
|
|
49
|
+
|
|
50
|
+
A. You can copyright only the *changes* or the *additions* you made to the
|
|
51
|
+
NCBI source code. You should identify unambiguously those sections of
|
|
52
|
+
the code that were modified, e.g. by commenting any changes you made
|
|
53
|
+
in the code you distribute. Therefore, your license has to make clear
|
|
54
|
+
to users that your product is a combination of code that is public domain
|
|
55
|
+
within the U.S. (but may be subject to copyright by the U.S. in foreign
|
|
56
|
+
countries) and code that has been created or modified by you.
|
|
57
|
+
|
|
58
|
+
--------------------------------------------------------------
|
|
59
|
+
Q. Can we (re)license all or part of the USDA - ARS source code?
|
|
60
|
+
|
|
61
|
+
A. No, you cannot license or relicense the source code written by USDA - ARS
|
|
62
|
+
since you cannot claim any copyright in the software that was developed
|
|
63
|
+
at USDA - ARS as a 'government work' and consequently is in the public
|
|
64
|
+
domain within the U.S.
|
|
65
|
+
|
|
66
|
+
--------------------------------------------------------------
|
|
67
|
+
Q. What if these copyright guidelines are not clear enough or are not
|
|
68
|
+
applicable to my particular case?
|
|
69
|
+
|
|
70
|
+
A. Contact us. Send your questions to 'answers@usda.gov'.
|
|
71
|
+
--------------------------------------------------------------
|
|
72
|
+
|
|
73
|
+
This file was modified from the NCBI Boilerplate LICENSE file
|
|
@@ -0,0 +1,261 @@
|
|
|
1
|
+
Metadata-Version: 2.4
|
|
2
|
+
Name: fair-bioheaders
|
|
3
|
+
Version: 0.4.0
|
|
4
|
+
Summary: Convert and validate FAIR-bioHeaders (FHR) metadata in JSON, YAML, FASTA, GFA, and HTML (formerly the fhr package)
|
|
5
|
+
License: USDA-ARS
|
|
6
|
+
License-File: LICENSE
|
|
7
|
+
Keywords: FAIR,FHR,genome,metadata,fasta,gfa
|
|
8
|
+
Author: David Molik
|
|
9
|
+
Author-email: david.molik@usda.gov
|
|
10
|
+
Requires-Python: >=3.9,<4.0
|
|
11
|
+
Classifier: License :: Other/Proprietary License
|
|
12
|
+
Classifier: Programming Language :: Python :: 3
|
|
13
|
+
Classifier: Programming Language :: Python :: 3.9
|
|
14
|
+
Classifier: Programming Language :: Python :: 3.10
|
|
15
|
+
Classifier: Programming Language :: Python :: 3.11
|
|
16
|
+
Classifier: Programming Language :: Python :: 3.12
|
|
17
|
+
Classifier: Programming Language :: Python :: 3.13
|
|
18
|
+
Classifier: Programming Language :: Python :: 3.14
|
|
19
|
+
Requires-Dist: jsonschema[format] (>=4.21.1,<5.0.0)
|
|
20
|
+
Requires-Dist: pyyaml (>=6.0.1,<7.0.0)
|
|
21
|
+
Project-URL: Bug Tracker, https://github.com/FAIR-bioHeaders/FAIR-bioHeaders-Tools/issues
|
|
22
|
+
Project-URL: Changelog, https://github.com/FAIR-bioHeaders/FAIR-bioHeaders-Tools/blob/main/CHANGELOG.md
|
|
23
|
+
Project-URL: Documentation, https://github.com/FAIR-bioHeaders/FHR-Specification
|
|
24
|
+
Project-URL: Homepage, https://github.com/FAIR-bioHeaders/FAIR-bioHeaders-Tools
|
|
25
|
+
Project-URL: Repository, https://github.com/FAIR-bioHeaders/FAIR-bioHeaders-Tools
|
|
26
|
+
Description-Content-Type: text/markdown
|
|
27
|
+
|
|
28
|
+
# FAIR-bioHeaders Tools
|
|
29
|
+
|
|
30
|
+
[](https://github.com/FAIR-bioHeaders/FAIR-bioHeaders-Tools/actions/workflows/pytest.yaml)
|
|
31
|
+
[](https://doi.org/10.5281/zenodo.6762549)
|
|
32
|
+
[](https://doi.org/10.5281/zenodo.6762547)
|
|
33
|
+
|
|
34
|
+
Convert and validate FAIR-bioHeaders Reference genome (FHR) metadata in JSON,
|
|
35
|
+
YAML, FASTA, GFA, and HTML microdata. See
|
|
36
|
+
[FHR-Specification](https://github.com/FAIR-bioHeaders/FHR-Specification) for the
|
|
37
|
+
schema and metadata design. This is the `fair-bioheaders` package, version
|
|
38
|
+
**0.4.0**, formerly the FHR File Converter (`fhr`); see
|
|
39
|
+
[Renamed from fhr](#renamed-from-fhr).
|
|
40
|
+
|
|
41
|
+
## Install
|
|
42
|
+
|
|
43
|
+
For a published release:
|
|
44
|
+
|
|
45
|
+
```bash
|
|
46
|
+
python -m pip install fair-bioheaders
|
|
47
|
+
```
|
|
48
|
+
|
|
49
|
+
For this checkout and its development checks (Python 3.9 or later):
|
|
50
|
+
|
|
51
|
+
```bash
|
|
52
|
+
python -m pip install poetry
|
|
53
|
+
poetry install
|
|
54
|
+
poetry run pytest
|
|
55
|
+
poetry run ruff check .
|
|
56
|
+
poetry run isort . --check-only
|
|
57
|
+
poetry run black . --check
|
|
58
|
+
```
|
|
59
|
+
|
|
60
|
+
## Commands
|
|
61
|
+
|
|
62
|
+
```bash
|
|
63
|
+
bioheaders convert examples/example.fhr.yaml /tmp/example.fhr.json
|
|
64
|
+
bioheaders validate /tmp/example.fhr.json
|
|
65
|
+
bioheaders combine examples/example.fhr.yaml genome.fasta -o genome.fhr.fasta
|
|
66
|
+
bioheaders verify genome.fhr.fasta
|
|
67
|
+
bioheaders strip genome.fhr.fasta genome.stripped.fasta
|
|
68
|
+
bioheaders combine examples/example.fhr.yaml assembly.gfa -o assembly.fhr.gfa
|
|
69
|
+
bioheaders verify assembly.fhr.gfa
|
|
70
|
+
bioheaders strip assembly.fhr.gfa assembly.stripped.gfa
|
|
71
|
+
```
|
|
72
|
+
|
|
73
|
+
`combine`, `strip`, and `verify` (also called `checksum`) take the FASTA or GFA
|
|
74
|
+
type from the sequence file's extension, ignoring `.gz` or `.bgz`; give
|
|
75
|
+
`--type fasta` or `--type gfa` for stdin or other extensions. `bioheaders
|
|
76
|
+
--version` prints the version and `bioheaders COMMAND --help` describes each
|
|
77
|
+
command. The original commands remain available with the same options and output:
|
|
78
|
+
|
|
79
|
+
| `bioheaders` command | Original command |
|
|
80
|
+
| --- | --- |
|
|
81
|
+
| `convert INPUT OUTPUT` | `fhr-convert` |
|
|
82
|
+
| `validate INPUT` | `fhr-validate` |
|
|
83
|
+
| `combine METADATA SEQUENCE` | `fhr-fasta-combine`, `fhr-gfa-combine` |
|
|
84
|
+
| `strip INPUT [OUTPUT]` | `fhr-fasta-strip`, `fhr-gfa-strip` |
|
|
85
|
+
| `verify INPUT` | `fhr-fasta-validate`, `fhr-gfa-validate` |
|
|
86
|
+
|
|
87
|
+
`convert INPUT OUTPUT` detects `.json`, `.yaml`/`.yml`, `.fasta`/`.fa`/`.fna`,
|
|
88
|
+
`.gfa`, and `.html` from extensions. It validates metadata before writing output;
|
|
89
|
+
FASTA/GFA output contains a metadata header only. To include sequence data, use
|
|
90
|
+
`combine METADATA SEQUENCE [-o OUTPUT]`, whose default output is `SEQUENCE.fhr.fasta`
|
|
91
|
+
or `SEQUENCE.fhr.gfa` with the last extension replaced. Existing FHR header lines
|
|
92
|
+
are replaced. Inputs are not overwritten by sequence helpers.
|
|
93
|
+
|
|
94
|
+
`strip INPUT [OUTPUT]` writes to stdout if OUTPUT is omitted. Only FHR-prefixed
|
|
95
|
+
lines are removed; other bytes, including ordinary comments and CRLF endings,
|
|
96
|
+
are preserved. `validate` checks metadata, while `verify` (`fhr-fasta-validate`,
|
|
97
|
+
`fhr-gfa-validate`) also verifies the exact-byte file checksum. Failures exit
|
|
98
|
+
with status 1.
|
|
99
|
+
|
|
100
|
+
FASTA/GFA commands stream their input in 1 MiB chunks, so memory use does not grow
|
|
101
|
+
with file size (about 35 MB peak for a 1 GB FASTA). Validate reads the file once;
|
|
102
|
+
combine reads it twice. FHR header lines are limited to 16 MiB in total. Outputs
|
|
103
|
+
are written to a temporary file in the destination directory and then renamed,
|
|
104
|
+
so a failed command leaves no partial output.
|
|
105
|
+
|
|
106
|
+
### Compressed files and pipes
|
|
107
|
+
|
|
108
|
+
```bash
|
|
109
|
+
bioheaders combine examples/example.fhr.yaml genome.fa.gz # writes genome.fhr.fasta.gz
|
|
110
|
+
bioheaders verify genome.fhr.fasta.gz
|
|
111
|
+
bioheaders strip genome.fhr.fasta.gz genome.stripped.fa.gz
|
|
112
|
+
zcat genome.fa.gz | bioheaders combine --type fasta examples/example.fhr.yaml - > genome.fhr.fasta
|
|
113
|
+
curl -sL https://example.org/genome.fhr.fasta.gz | bioheaders verify --type fasta -
|
|
114
|
+
bioheaders convert genome.fhr.fasta.gz - --to json
|
|
115
|
+
bioheaders convert - metadata.yaml --from json < metadata.json
|
|
116
|
+
```
|
|
117
|
+
|
|
118
|
+
Gzip input, including multi-member gzip and BGZF, is recognized by its magic
|
|
119
|
+
bytes rather than its extension and decompressed as it streams. The checksum
|
|
120
|
+
covers the decompressed FASTA/GFA bytes, so `genome.fa` and any gzip or BGZF
|
|
121
|
+
compression of it have the same checksum. A corrupt or truncated gzip file is an
|
|
122
|
+
error. Formats are taken from the extension before `.gz` or `.bgz`.
|
|
123
|
+
|
|
124
|
+
Outputs whose path ends in `.gz` or `.bgz` are written as BGZF, which `gzip`,
|
|
125
|
+
`zcat`, `bgzip`, and htslib read. A combine without `-o` keeps the input's
|
|
126
|
+
compression extension. Other outputs, and stdout, are never compressed. Note that
|
|
127
|
+
`samtools faidx` rejects FASTA files containing `;` lines, compressed or not, so
|
|
128
|
+
index a stripped copy for random access.
|
|
129
|
+
|
|
130
|
+
`-` reads stdin or writes stdout: the input of verify, strip, combine (the
|
|
131
|
+
sequence), convert, and validate, and the output of strip, convert, and
|
|
132
|
+
`combine -o`. Combine of stdin writes stdout by default. Convert and validate
|
|
133
|
+
need `--from FORMAT` for stdin and convert needs `--to FORMAT` for stdout
|
|
134
|
+
(`json`, `yaml`, `fasta`, `gfa`, or `html`); the `bioheaders` sequence commands
|
|
135
|
+
need `--type` for stdin. Stdin is read
|
|
136
|
+
once: combine spools the stripped sequence to a temporary file in `TMPDIR` (as
|
|
137
|
+
large as the sequence) while hashing it. Strip from stdin to stdout writes as it
|
|
138
|
+
reads, so an error late in the input can follow partial output; check the exit
|
|
139
|
+
status. Strip of a file to stdout still checks the whole file before writing.
|
|
140
|
+
|
|
141
|
+
## Python
|
|
142
|
+
|
|
143
|
+
```python
|
|
144
|
+
from bioheaders import fhr
|
|
145
|
+
|
|
146
|
+
metadata = fhr()
|
|
147
|
+
with open("examples/example.fhr.yaml", encoding="utf-8") as stream:
|
|
148
|
+
metadata.input_yaml(stream)
|
|
149
|
+
metadata.fhr_validate()
|
|
150
|
+
print(metadata.output_json())
|
|
151
|
+
```
|
|
152
|
+
|
|
153
|
+
Input methods accept text, UTF-8 bytes, or readable streams. Optional fields stay
|
|
154
|
+
absent until supplied. The mapping is preserved across supported format round
|
|
155
|
+
trips; HTML values are escaped and explicitly typed. Keywords can initialize an
|
|
156
|
+
instance (`fhr(genome="example")`), and fields remain accessible as attributes.
|
|
157
|
+
The installed schema is loaded from package resources rather than the working
|
|
158
|
+
directory, so commands work outside the checkout. No schema is downloaded at runtime.
|
|
159
|
+
`bioheaders.cli` holds the command implementations and the byte-level helpers
|
|
160
|
+
(`checksum`, `combine`, `strip_header`, `open_input`, `write_to`, and others).
|
|
161
|
+
|
|
162
|
+
## Renamed from fhr
|
|
163
|
+
|
|
164
|
+
In 0.4.0 the PyPI package `fhr` was renamed `fair-bioheaders`, the Python package
|
|
165
|
+
`fhr` was renamed `bioheaders`, and the repository moved from FHR-File-Converter
|
|
166
|
+
to [FAIR-bioHeaders-Tools](https://github.com/FAIR-bioHeaders/FAIR-bioHeaders-Tools)
|
|
167
|
+
(old links redirect). The new names leave room for other FAIR-bioHeaders header
|
|
168
|
+
types. Nothing that worked with 0.3 stops working:
|
|
169
|
+
|
|
170
|
+
- The eight `fhr-*` commands are unchanged and print no deprecation notices, so
|
|
171
|
+
scripts and pipelines that read their output or stderr keep working. They are
|
|
172
|
+
not deprecated, but new scripts can use `bioheaders`.
|
|
173
|
+
- `import fhr`, `from fhr import fhr`, and `from fhr.cli import ...` still work:
|
|
174
|
+
`fhr` and `fhr.cli` are the same module objects as `bioheaders` and
|
|
175
|
+
`bioheaders.cli`. Importing `fhr` emits one `DeprecationWarning`; replace
|
|
176
|
+
`fhr` with `bioheaders` in imports. The `fhr` metadata class keeps its name.
|
|
177
|
+
- `pip install fhr` installs `fhr` 0.4.0, a compatibility package without code
|
|
178
|
+
that requires the same version of `fair-bioheaders`, so existing requirements
|
|
179
|
+
and `pip install -U fhr` keep working. Prefer `fair-bioheaders` in new
|
|
180
|
+
requirements. The compatibility package also installs the `fhr-*` commands, so
|
|
181
|
+
if you uninstall `fhr` 0.4 later, restore them with
|
|
182
|
+
`python -m pip install --force-reinstall --no-deps fair-bioheaders`.
|
|
183
|
+
- In an environment that still has `fhr` 0.3, run `python -m pip uninstall fhr`
|
|
184
|
+
before `python -m pip install fair-bioheaders` (or upgrade with
|
|
185
|
+
`python -m pip install -U fhr`); otherwise the old `fhr` package files shadow
|
|
186
|
+
the new compatibility module.
|
|
187
|
+
- Citations and Zenodo DOIs are unchanged.
|
|
188
|
+
|
|
189
|
+
Checksum helpers require SHA-512/256 support in the Python build. Some Apple
|
|
190
|
+
system Python builds omit it; use an OpenSSL-enabled Python distribution.
|
|
191
|
+
Metadata conversion and validation do not require that hash implementation.
|
|
192
|
+
|
|
193
|
+
## v0.3 compatibility and identity
|
|
194
|
+
|
|
195
|
+
- Required fields and `schemaVersion: 1` remain unchanged since v0.3.0.
|
|
196
|
+
- `assemblySoftware` accepts a legacy string or optional structured software
|
|
197
|
+
objects with name, URI, version, and command options. `assemblyProtocol` is a URI.
|
|
198
|
+
- `vitalStats.N90` is base pairs; `vitalStats.gcContent` is 0–100 percent.
|
|
199
|
+
- Optional `seqcol_id` stores a supplied 32-character base64url top-level refget
|
|
200
|
+
digest. The converter preserves it; it does not compute or verify SeqCol identity.
|
|
201
|
+
- Checksums use base64 SHA-512/256 of all original file bytes except the one scalar
|
|
202
|
+
checksum header line, including its newline. Metadata is covered. MD5 hex and
|
|
203
|
+
payload-only checksums from old examples are not valid v0.3 checksums; recombine
|
|
204
|
+
metadata with the original sequence file to calculate the new value.
|
|
205
|
+
- The checksum value must be a single-line scalar on the checksum line. Header
|
|
206
|
+
lines must be UTF-8 and must not contain U+0085, U+2028, or U+2029. FASTA/GFA
|
|
207
|
+
files must not begin with a UTF-8 byte order mark. Other sequence bytes may use
|
|
208
|
+
any encoding.
|
|
209
|
+
- Metadata must be JSON-compatible: duplicate keys, YAML anchors, aliases, and
|
|
210
|
+
merge keys are rejected.
|
|
211
|
+
- `;~`/`#~` lines must form the leading header block, before the first FASTA `>`
|
|
212
|
+
line or the first GFA record line; ordinary comments and blank lines may be
|
|
213
|
+
mixed in. A `;~`/`#~` line after sequence data, including a concatenated
|
|
214
|
+
second file, is an error.
|
|
215
|
+
- HTML exports use FHR item scopes and `data-fhr-type` annotations for lossless
|
|
216
|
+
arrays, numbers, objects, and strings. External microdata must represent nested
|
|
217
|
+
items properly; incomplete legacy markup may need regeneration.
|
|
218
|
+
- Incomplete constructors now emit missing fields rather than empty defaults.
|
|
219
|
+
Validate after loading to obtain actionable schema errors. Obsolete positional
|
|
220
|
+
constructor arguments should be converted to keywords.
|
|
221
|
+
|
|
222
|
+
See the [format reference](https://github.com/FAIR-bioHeaders/FHR-Specification/blob/main/docs/FORMAT.md)
|
|
223
|
+
and [release notes](CHANGELOG.md). JSON/YAML/HTML example identifiers are synthetic.
|
|
224
|
+
The FASTA/GFA fixtures contain verified FHR checksums; their SeqCol IDs are placeholders.
|
|
225
|
+
|
|
226
|
+
## Docker
|
|
227
|
+
|
|
228
|
+
```bash
|
|
229
|
+
docker build -t fair-bioheaders-tools .
|
|
230
|
+
docker run --rm fair-bioheaders-tools --help
|
|
231
|
+
```
|
|
232
|
+
|
|
233
|
+
The image runs `fhr-convert`; use `--entrypoint bioheaders` for the other
|
|
234
|
+
commands. Mount inputs/outputs into the container for conversion. See
|
|
235
|
+
[CONTRIBUTING](CONTRIBUTING.md), [CODE_OF_CONDUCT](CODE_OF_CONDUCT.md),
|
|
236
|
+
[SECURITY](SECURITY.md), and [AGENTS](AGENTS.md) for project guidance.
|
|
237
|
+
|
|
238
|
+
## Citing FHR
|
|
239
|
+
|
|
240
|
+
Chicago bibliography entries are used below. Cite the published paper for a
|
|
241
|
+
general description of FHR; cite the specification or converter when using that
|
|
242
|
+
resource directly. The software and specification links are concept DOIs; for a
|
|
243
|
+
specific release, use the corresponding version DOI from Zenodo. Authors and
|
|
244
|
+
years follow the records resolved by the concept DOIs at the v0.3 documentation
|
|
245
|
+
update, and can change as later records are published.
|
|
246
|
+
|
|
247
|
+
### Published paper
|
|
248
|
+
|
|
249
|
+
Wright, Adam, Mark D. Wilkinson, Christopher Mungall, Scott Cain, Stephen Richards, Paul Sternberg, Ellen Provin, Jonathan L. Jacobs, Scott Geib, Daniela Raciti, Karen Yook, Lincoln Stein, and David C. Molik. “FAIR Header Reference Genome: A TRUSTworthy Standard.” *Briefings in Bioinformatics* 25, no. 3 (2024): bbae122. https://doi.org/10.1093/bib/bbae122.
|
|
250
|
+
|
|
251
|
+
### Specification
|
|
252
|
+
|
|
253
|
+
Molik, David. *FHR Specification*. Data set. 2022. https://doi.org/10.5281/zenodo.6762549.
|
|
254
|
+
|
|
255
|
+
### Converter
|
|
256
|
+
|
|
257
|
+
Molik, David, and Adam Wright. *FHR File Converter*. Computer software. 2024. https://doi.org/10.5281/zenodo.6762547.
|
|
258
|
+
|
|
259
|
+
Machine-readable entries are maintained in
|
|
260
|
+
[FHR-Citation](https://github.com/FAIR-bioHeaders/FHR-Citation/blob/main/citation.bib).
|
|
261
|
+
|
|
@@ -0,0 +1,233 @@
|
|
|
1
|
+
# FAIR-bioHeaders Tools
|
|
2
|
+
|
|
3
|
+
[](https://github.com/FAIR-bioHeaders/FAIR-bioHeaders-Tools/actions/workflows/pytest.yaml)
|
|
4
|
+
[](https://doi.org/10.5281/zenodo.6762549)
|
|
5
|
+
[](https://doi.org/10.5281/zenodo.6762547)
|
|
6
|
+
|
|
7
|
+
Convert and validate FAIR-bioHeaders Reference genome (FHR) metadata in JSON,
|
|
8
|
+
YAML, FASTA, GFA, and HTML microdata. See
|
|
9
|
+
[FHR-Specification](https://github.com/FAIR-bioHeaders/FHR-Specification) for the
|
|
10
|
+
schema and metadata design. This is the `fair-bioheaders` package, version
|
|
11
|
+
**0.4.0**, formerly the FHR File Converter (`fhr`); see
|
|
12
|
+
[Renamed from fhr](#renamed-from-fhr).
|
|
13
|
+
|
|
14
|
+
## Install
|
|
15
|
+
|
|
16
|
+
For a published release:
|
|
17
|
+
|
|
18
|
+
```bash
|
|
19
|
+
python -m pip install fair-bioheaders
|
|
20
|
+
```
|
|
21
|
+
|
|
22
|
+
For this checkout and its development checks (Python 3.9 or later):
|
|
23
|
+
|
|
24
|
+
```bash
|
|
25
|
+
python -m pip install poetry
|
|
26
|
+
poetry install
|
|
27
|
+
poetry run pytest
|
|
28
|
+
poetry run ruff check .
|
|
29
|
+
poetry run isort . --check-only
|
|
30
|
+
poetry run black . --check
|
|
31
|
+
```
|
|
32
|
+
|
|
33
|
+
## Commands
|
|
34
|
+
|
|
35
|
+
```bash
|
|
36
|
+
bioheaders convert examples/example.fhr.yaml /tmp/example.fhr.json
|
|
37
|
+
bioheaders validate /tmp/example.fhr.json
|
|
38
|
+
bioheaders combine examples/example.fhr.yaml genome.fasta -o genome.fhr.fasta
|
|
39
|
+
bioheaders verify genome.fhr.fasta
|
|
40
|
+
bioheaders strip genome.fhr.fasta genome.stripped.fasta
|
|
41
|
+
bioheaders combine examples/example.fhr.yaml assembly.gfa -o assembly.fhr.gfa
|
|
42
|
+
bioheaders verify assembly.fhr.gfa
|
|
43
|
+
bioheaders strip assembly.fhr.gfa assembly.stripped.gfa
|
|
44
|
+
```
|
|
45
|
+
|
|
46
|
+
`combine`, `strip`, and `verify` (also called `checksum`) take the FASTA or GFA
|
|
47
|
+
type from the sequence file's extension, ignoring `.gz` or `.bgz`; give
|
|
48
|
+
`--type fasta` or `--type gfa` for stdin or other extensions. `bioheaders
|
|
49
|
+
--version` prints the version and `bioheaders COMMAND --help` describes each
|
|
50
|
+
command. The original commands remain available with the same options and output:
|
|
51
|
+
|
|
52
|
+
| `bioheaders` command | Original command |
|
|
53
|
+
| --- | --- |
|
|
54
|
+
| `convert INPUT OUTPUT` | `fhr-convert` |
|
|
55
|
+
| `validate INPUT` | `fhr-validate` |
|
|
56
|
+
| `combine METADATA SEQUENCE` | `fhr-fasta-combine`, `fhr-gfa-combine` |
|
|
57
|
+
| `strip INPUT [OUTPUT]` | `fhr-fasta-strip`, `fhr-gfa-strip` |
|
|
58
|
+
| `verify INPUT` | `fhr-fasta-validate`, `fhr-gfa-validate` |
|
|
59
|
+
|
|
60
|
+
`convert INPUT OUTPUT` detects `.json`, `.yaml`/`.yml`, `.fasta`/`.fa`/`.fna`,
|
|
61
|
+
`.gfa`, and `.html` from extensions. It validates metadata before writing output;
|
|
62
|
+
FASTA/GFA output contains a metadata header only. To include sequence data, use
|
|
63
|
+
`combine METADATA SEQUENCE [-o OUTPUT]`, whose default output is `SEQUENCE.fhr.fasta`
|
|
64
|
+
or `SEQUENCE.fhr.gfa` with the last extension replaced. Existing FHR header lines
|
|
65
|
+
are replaced. Inputs are not overwritten by sequence helpers.
|
|
66
|
+
|
|
67
|
+
`strip INPUT [OUTPUT]` writes to stdout if OUTPUT is omitted. Only FHR-prefixed
|
|
68
|
+
lines are removed; other bytes, including ordinary comments and CRLF endings,
|
|
69
|
+
are preserved. `validate` checks metadata, while `verify` (`fhr-fasta-validate`,
|
|
70
|
+
`fhr-gfa-validate`) also verifies the exact-byte file checksum. Failures exit
|
|
71
|
+
with status 1.
|
|
72
|
+
|
|
73
|
+
FASTA/GFA commands stream their input in 1 MiB chunks, so memory use does not grow
|
|
74
|
+
with file size (about 35 MB peak for a 1 GB FASTA). Validate reads the file once;
|
|
75
|
+
combine reads it twice. FHR header lines are limited to 16 MiB in total. Outputs
|
|
76
|
+
are written to a temporary file in the destination directory and then renamed,
|
|
77
|
+
so a failed command leaves no partial output.
|
|
78
|
+
|
|
79
|
+
### Compressed files and pipes
|
|
80
|
+
|
|
81
|
+
```bash
|
|
82
|
+
bioheaders combine examples/example.fhr.yaml genome.fa.gz # writes genome.fhr.fasta.gz
|
|
83
|
+
bioheaders verify genome.fhr.fasta.gz
|
|
84
|
+
bioheaders strip genome.fhr.fasta.gz genome.stripped.fa.gz
|
|
85
|
+
zcat genome.fa.gz | bioheaders combine --type fasta examples/example.fhr.yaml - > genome.fhr.fasta
|
|
86
|
+
curl -sL https://example.org/genome.fhr.fasta.gz | bioheaders verify --type fasta -
|
|
87
|
+
bioheaders convert genome.fhr.fasta.gz - --to json
|
|
88
|
+
bioheaders convert - metadata.yaml --from json < metadata.json
|
|
89
|
+
```
|
|
90
|
+
|
|
91
|
+
Gzip input, including multi-member gzip and BGZF, is recognized by its magic
|
|
92
|
+
bytes rather than its extension and decompressed as it streams. The checksum
|
|
93
|
+
covers the decompressed FASTA/GFA bytes, so `genome.fa` and any gzip or BGZF
|
|
94
|
+
compression of it have the same checksum. A corrupt or truncated gzip file is an
|
|
95
|
+
error. Formats are taken from the extension before `.gz` or `.bgz`.
|
|
96
|
+
|
|
97
|
+
Outputs whose path ends in `.gz` or `.bgz` are written as BGZF, which `gzip`,
|
|
98
|
+
`zcat`, `bgzip`, and htslib read. A combine without `-o` keeps the input's
|
|
99
|
+
compression extension. Other outputs, and stdout, are never compressed. Note that
|
|
100
|
+
`samtools faidx` rejects FASTA files containing `;` lines, compressed or not, so
|
|
101
|
+
index a stripped copy for random access.
|
|
102
|
+
|
|
103
|
+
`-` reads stdin or writes stdout: the input of verify, strip, combine (the
|
|
104
|
+
sequence), convert, and validate, and the output of strip, convert, and
|
|
105
|
+
`combine -o`. Combine of stdin writes stdout by default. Convert and validate
|
|
106
|
+
need `--from FORMAT` for stdin and convert needs `--to FORMAT` for stdout
|
|
107
|
+
(`json`, `yaml`, `fasta`, `gfa`, or `html`); the `bioheaders` sequence commands
|
|
108
|
+
need `--type` for stdin. Stdin is read
|
|
109
|
+
once: combine spools the stripped sequence to a temporary file in `TMPDIR` (as
|
|
110
|
+
large as the sequence) while hashing it. Strip from stdin to stdout writes as it
|
|
111
|
+
reads, so an error late in the input can follow partial output; check the exit
|
|
112
|
+
status. Strip of a file to stdout still checks the whole file before writing.
|
|
113
|
+
|
|
114
|
+
## Python
|
|
115
|
+
|
|
116
|
+
```python
|
|
117
|
+
from bioheaders import fhr
|
|
118
|
+
|
|
119
|
+
metadata = fhr()
|
|
120
|
+
with open("examples/example.fhr.yaml", encoding="utf-8") as stream:
|
|
121
|
+
metadata.input_yaml(stream)
|
|
122
|
+
metadata.fhr_validate()
|
|
123
|
+
print(metadata.output_json())
|
|
124
|
+
```
|
|
125
|
+
|
|
126
|
+
Input methods accept text, UTF-8 bytes, or readable streams. Optional fields stay
|
|
127
|
+
absent until supplied. The mapping is preserved across supported format round
|
|
128
|
+
trips; HTML values are escaped and explicitly typed. Keywords can initialize an
|
|
129
|
+
instance (`fhr(genome="example")`), and fields remain accessible as attributes.
|
|
130
|
+
The installed schema is loaded from package resources rather than the working
|
|
131
|
+
directory, so commands work outside the checkout. No schema is downloaded at runtime.
|
|
132
|
+
`bioheaders.cli` holds the command implementations and the byte-level helpers
|
|
133
|
+
(`checksum`, `combine`, `strip_header`, `open_input`, `write_to`, and others).
|
|
134
|
+
|
|
135
|
+
## Renamed from fhr
|
|
136
|
+
|
|
137
|
+
In 0.4.0 the PyPI package `fhr` was renamed `fair-bioheaders`, the Python package
|
|
138
|
+
`fhr` was renamed `bioheaders`, and the repository moved from FHR-File-Converter
|
|
139
|
+
to [FAIR-bioHeaders-Tools](https://github.com/FAIR-bioHeaders/FAIR-bioHeaders-Tools)
|
|
140
|
+
(old links redirect). The new names leave room for other FAIR-bioHeaders header
|
|
141
|
+
types. Nothing that worked with 0.3 stops working:
|
|
142
|
+
|
|
143
|
+
- The eight `fhr-*` commands are unchanged and print no deprecation notices, so
|
|
144
|
+
scripts and pipelines that read their output or stderr keep working. They are
|
|
145
|
+
not deprecated, but new scripts can use `bioheaders`.
|
|
146
|
+
- `import fhr`, `from fhr import fhr`, and `from fhr.cli import ...` still work:
|
|
147
|
+
`fhr` and `fhr.cli` are the same module objects as `bioheaders` and
|
|
148
|
+
`bioheaders.cli`. Importing `fhr` emits one `DeprecationWarning`; replace
|
|
149
|
+
`fhr` with `bioheaders` in imports. The `fhr` metadata class keeps its name.
|
|
150
|
+
- `pip install fhr` installs `fhr` 0.4.0, a compatibility package without code
|
|
151
|
+
that requires the same version of `fair-bioheaders`, so existing requirements
|
|
152
|
+
and `pip install -U fhr` keep working. Prefer `fair-bioheaders` in new
|
|
153
|
+
requirements. The compatibility package also installs the `fhr-*` commands, so
|
|
154
|
+
if you uninstall `fhr` 0.4 later, restore them with
|
|
155
|
+
`python -m pip install --force-reinstall --no-deps fair-bioheaders`.
|
|
156
|
+
- In an environment that still has `fhr` 0.3, run `python -m pip uninstall fhr`
|
|
157
|
+
before `python -m pip install fair-bioheaders` (or upgrade with
|
|
158
|
+
`python -m pip install -U fhr`); otherwise the old `fhr` package files shadow
|
|
159
|
+
the new compatibility module.
|
|
160
|
+
- Citations and Zenodo DOIs are unchanged.
|
|
161
|
+
|
|
162
|
+
Checksum helpers require SHA-512/256 support in the Python build. Some Apple
|
|
163
|
+
system Python builds omit it; use an OpenSSL-enabled Python distribution.
|
|
164
|
+
Metadata conversion and validation do not require that hash implementation.
|
|
165
|
+
|
|
166
|
+
## v0.3 compatibility and identity
|
|
167
|
+
|
|
168
|
+
- Required fields and `schemaVersion: 1` remain unchanged since v0.3.0.
|
|
169
|
+
- `assemblySoftware` accepts a legacy string or optional structured software
|
|
170
|
+
objects with name, URI, version, and command options. `assemblyProtocol` is a URI.
|
|
171
|
+
- `vitalStats.N90` is base pairs; `vitalStats.gcContent` is 0–100 percent.
|
|
172
|
+
- Optional `seqcol_id` stores a supplied 32-character base64url top-level refget
|
|
173
|
+
digest. The converter preserves it; it does not compute or verify SeqCol identity.
|
|
174
|
+
- Checksums use base64 SHA-512/256 of all original file bytes except the one scalar
|
|
175
|
+
checksum header line, including its newline. Metadata is covered. MD5 hex and
|
|
176
|
+
payload-only checksums from old examples are not valid v0.3 checksums; recombine
|
|
177
|
+
metadata with the original sequence file to calculate the new value.
|
|
178
|
+
- The checksum value must be a single-line scalar on the checksum line. Header
|
|
179
|
+
lines must be UTF-8 and must not contain U+0085, U+2028, or U+2029. FASTA/GFA
|
|
180
|
+
files must not begin with a UTF-8 byte order mark. Other sequence bytes may use
|
|
181
|
+
any encoding.
|
|
182
|
+
- Metadata must be JSON-compatible: duplicate keys, YAML anchors, aliases, and
|
|
183
|
+
merge keys are rejected.
|
|
184
|
+
- `;~`/`#~` lines must form the leading header block, before the first FASTA `>`
|
|
185
|
+
line or the first GFA record line; ordinary comments and blank lines may be
|
|
186
|
+
mixed in. A `;~`/`#~` line after sequence data, including a concatenated
|
|
187
|
+
second file, is an error.
|
|
188
|
+
- HTML exports use FHR item scopes and `data-fhr-type` annotations for lossless
|
|
189
|
+
arrays, numbers, objects, and strings. External microdata must represent nested
|
|
190
|
+
items properly; incomplete legacy markup may need regeneration.
|
|
191
|
+
- Incomplete constructors now emit missing fields rather than empty defaults.
|
|
192
|
+
Validate after loading to obtain actionable schema errors. Obsolete positional
|
|
193
|
+
constructor arguments should be converted to keywords.
|
|
194
|
+
|
|
195
|
+
See the [format reference](https://github.com/FAIR-bioHeaders/FHR-Specification/blob/main/docs/FORMAT.md)
|
|
196
|
+
and [release notes](CHANGELOG.md). JSON/YAML/HTML example identifiers are synthetic.
|
|
197
|
+
The FASTA/GFA fixtures contain verified FHR checksums; their SeqCol IDs are placeholders.
|
|
198
|
+
|
|
199
|
+
## Docker
|
|
200
|
+
|
|
201
|
+
```bash
|
|
202
|
+
docker build -t fair-bioheaders-tools .
|
|
203
|
+
docker run --rm fair-bioheaders-tools --help
|
|
204
|
+
```
|
|
205
|
+
|
|
206
|
+
The image runs `fhr-convert`; use `--entrypoint bioheaders` for the other
|
|
207
|
+
commands. Mount inputs/outputs into the container for conversion. See
|
|
208
|
+
[CONTRIBUTING](CONTRIBUTING.md), [CODE_OF_CONDUCT](CODE_OF_CONDUCT.md),
|
|
209
|
+
[SECURITY](SECURITY.md), and [AGENTS](AGENTS.md) for project guidance.
|
|
210
|
+
|
|
211
|
+
## Citing FHR
|
|
212
|
+
|
|
213
|
+
Chicago bibliography entries are used below. Cite the published paper for a
|
|
214
|
+
general description of FHR; cite the specification or converter when using that
|
|
215
|
+
resource directly. The software and specification links are concept DOIs; for a
|
|
216
|
+
specific release, use the corresponding version DOI from Zenodo. Authors and
|
|
217
|
+
years follow the records resolved by the concept DOIs at the v0.3 documentation
|
|
218
|
+
update, and can change as later records are published.
|
|
219
|
+
|
|
220
|
+
### Published paper
|
|
221
|
+
|
|
222
|
+
Wright, Adam, Mark D. Wilkinson, Christopher Mungall, Scott Cain, Stephen Richards, Paul Sternberg, Ellen Provin, Jonathan L. Jacobs, Scott Geib, Daniela Raciti, Karen Yook, Lincoln Stein, and David C. Molik. “FAIR Header Reference Genome: A TRUSTworthy Standard.” *Briefings in Bioinformatics* 25, no. 3 (2024): bbae122. https://doi.org/10.1093/bib/bbae122.
|
|
223
|
+
|
|
224
|
+
### Specification
|
|
225
|
+
|
|
226
|
+
Molik, David. *FHR Specification*. Data set. 2022. https://doi.org/10.5281/zenodo.6762549.
|
|
227
|
+
|
|
228
|
+
### Converter
|
|
229
|
+
|
|
230
|
+
Molik, David, and Adam Wright. *FHR File Converter*. Computer software. 2024. https://doi.org/10.5281/zenodo.6762547.
|
|
231
|
+
|
|
232
|
+
Machine-readable entries are maintained in
|
|
233
|
+
[FHR-Citation](https://github.com/FAIR-bioHeaders/FHR-Citation/blob/main/citation.bib).
|