py-ard 2.4.1__py3-none-any.whl
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- py_ard-2.4.1.dist-info/METADATA +784 -0
- py_ard-2.4.1.dist-info/RECORD +58 -0
- py_ard-2.4.1.dist-info/WHEEL +4 -0
- py_ard-2.4.1.dist-info/entry_points.txt +5 -0
- py_ard-2.4.1.dist-info/licenses/LICENSE +165 -0
- pyard/__init__.py +73 -0
- pyard/ard.py +487 -0
- pyard/blender.py +298 -0
- pyard/cli/__init__.py +22 -0
- pyard/cli/import_db.py +164 -0
- pyard/cli/reduce_csv.py +524 -0
- pyard/cli/redux.py +256 -0
- pyard/cli/status.py +117 -0
- pyard/config.py +115 -0
- pyard/constants.py +59 -0
- pyard/data_repository.py +546 -0
- pyard/db.py +718 -0
- pyard/dna_relshp.csv +36 -0
- pyard/drbx.py +65 -0
- pyard/exceptions.py +53 -0
- pyard/handlers/__init__.py +19 -0
- pyard/handlers/allele_handler.py +71 -0
- pyard/handlers/gl_string_processor.py +174 -0
- pyard/handlers/hats_handler.py +31 -0
- pyard/handlers/mac_handler.py +179 -0
- pyard/handlers/serology_handler.py +116 -0
- pyard/handlers/shortnull_handler.py +57 -0
- pyard/handlers/v2_handler.py +148 -0
- pyard/handlers/xx_handler.py +91 -0
- pyard/loader/CWD2.csv +1064 -0
- pyard/loader/__init__.py +7 -0
- pyard/loader/allele_list.py +63 -0
- pyard/loader/cwd.py +14 -0
- pyard/loader/g_group.py +71 -0
- pyard/loader/mac_codes.py +75 -0
- pyard/loader/p_group.py +78 -0
- pyard/loader/serology.py +161 -0
- pyard/loader/version.py +24 -0
- pyard/loader/wmda-data-model.md +378 -0
- pyard/mappings.py +53 -0
- pyard/misc.py +319 -0
- pyard/rc.py +50 -0
- pyard/reducers/__init__.py +28 -0
- pyard/reducers/base_reducer.py +113 -0
- pyard/reducers/default_reducer.py +100 -0
- pyard/reducers/exon_reducer.py +102 -0
- pyard/reducers/first_field_reducer.py +18 -0
- pyard/reducers/g_reducer.py +72 -0
- pyard/reducers/hats_reducer.py +43 -0
- pyard/reducers/lg_reducer.py +133 -0
- pyard/reducers/p_reducer.py +73 -0
- pyard/reducers/reducer_factory.py +48 -0
- pyard/reducers/s_reducer.py +111 -0
- pyard/reducers/u2_reducer.py +88 -0
- pyard/reducers/w_reducer.py +89 -0
- pyard/serology.py +413 -0
- pyard/simple_table.py +338 -0
- pyard/smart_sort.py +187 -0
|
@@ -0,0 +1,784 @@
|
|
|
1
|
+
Metadata-Version: 2.5
|
|
2
|
+
Name: py-ard
|
|
3
|
+
Version: 2.4.1
|
|
4
|
+
Summary: ARD reduction for HLA with Python
|
|
5
|
+
Project-URL: Homepage, https://github.com/nmdp-bioinformatics/py-ard
|
|
6
|
+
Project-URL: Repository, https://github.com/nmdp-bioinformatics/py-ard
|
|
7
|
+
Author-email: CIBMTR <cibmtr-pypi@nmdp.org>
|
|
8
|
+
License-Expression: LGPL-3.0-or-later
|
|
9
|
+
License-File: LICENSE
|
|
10
|
+
Keywords: pyard
|
|
11
|
+
Classifier: Development Status :: 5 - Production/Stable
|
|
12
|
+
Classifier: Intended Audience :: Science/Research
|
|
13
|
+
Classifier: Natural Language :: English
|
|
14
|
+
Classifier: Programming Language :: Python :: 3.8
|
|
15
|
+
Classifier: Programming Language :: Python :: 3.9
|
|
16
|
+
Classifier: Programming Language :: Python :: 3.10
|
|
17
|
+
Classifier: Programming Language :: Python :: 3.11
|
|
18
|
+
Classifier: Programming Language :: Python :: 3.12
|
|
19
|
+
Classifier: Programming Language :: Python :: 3.13
|
|
20
|
+
Classifier: Programming Language :: Python :: 3.14
|
|
21
|
+
Classifier: Topic :: Scientific/Engineering :: Bio-Informatics
|
|
22
|
+
Requires-Python: >=3.8
|
|
23
|
+
Requires-Dist: toml==0.10.2
|
|
24
|
+
Provides-Extra: deploy
|
|
25
|
+
Requires-Dist: connexion[flask,swagger-ui,uvicorn]==3.3.0; (python_version >= '3.10') and extra == 'deploy'
|
|
26
|
+
Requires-Dist: flask>=3.1.0; (python_version >= '3.10') and extra == 'deploy'
|
|
27
|
+
Requires-Dist: gunicorn==25.3.0; (python_version >= '3.10') and extra == 'deploy'
|
|
28
|
+
Requires-Dist: uvicorn>=0.45.0; (python_version >= '3.10') and extra == 'deploy'
|
|
29
|
+
Provides-Extra: script
|
|
30
|
+
Requires-Dist: pandas>=2.0; extra == 'script'
|
|
31
|
+
Description-Content-Type: text/markdown
|
|
32
|
+
|
|
33
|
+
# py-ard
|
|
34
|
+
|
|
35
|
+
Swiss army knife of **HLA** Nomenclature
|
|
36
|
+
|
|
37
|
+
[](https://pypi.python.org/pypi/py-ard)
|
|
38
|
+
|
|
39
|
+

|
|
40
|
+
|
|
41
|
+
**Note:**
|
|
42
|
+
|
|
43
|
+
- With `py-ard>=2.0.0`, the dependency on Pandas library has been removed.
|
|
44
|
+
|
|
45
|
+
---
|
|
46
|
+
|
|
47
|
+
### `py-ard` is ARD reduction for HLA in Python
|
|
48
|
+
|
|
49
|
+
Human leukocyte antigen (HLA) genes encode cell surface proteins that are important for immune regulation. Exons
|
|
50
|
+
encoding the Antigen Recognition Domain (ARD) are the most polymorphic region of HLA genes and are important for
|
|
51
|
+
donor/recipient [HLA matching](https://bethematch.org/patients-and-families/before-transplant/find-a-donor/hla-matching/).
|
|
52
|
+
The history of allele typing methods has played a major role in determining resolution and ambiguity of reported HLA
|
|
53
|
+
values. Although
|
|
54
|
+
HLA [nomenclature](https://www.theatlantic.com/magazine/archive/2023/04/clint-smith-nomenclature-poem/673097/) has not
|
|
55
|
+
always conformed to the same standard, it is now defined
|
|
56
|
+
by [The WHO Nomenclature Committee for Factors of the HLA System](https://hla.alleles.org/nomenclature/committee.html). `py-ard`
|
|
57
|
+
is aware of the variation in historical resolutions and grouping and is able to translate from one representation to
|
|
58
|
+
another based on alleles published quarterly by [IPD/IMGT-HLA](https://github.com/ANHIG/IMGTHLA/).
|
|
59
|
+
|
|
60
|
+
## Table of Contents
|
|
61
|
+
|
|
62
|
+
1. [Installation](#installation)
|
|
63
|
+
* [Install From PyPi](#install-from-pypi)
|
|
64
|
+
* [Install With Homebrew](#install-with-homebrew)
|
|
65
|
+
* [Install From Source](#install-from-source)
|
|
66
|
+
2. [Using `py-ard`](#using-py-ard)
|
|
67
|
+
* [Using `py-ard` from Python](#using-py-ard-from-python-code)
|
|
68
|
+
* [Using `py-ard` from R](#using-py-ard-from-r-code)
|
|
69
|
+
* [`.pyardrc` Configuration File](#pyardrc-configuration-file)
|
|
70
|
+
* [Perform Reduction](#reduce-typings)
|
|
71
|
+
* [DRBX blending](#perform-drb1-blending-with-drb3-drb4-and-drb5)
|
|
72
|
+
* [Expand/Lookup MAC](#mac-codes)
|
|
73
|
+
3. [Command Line Tools](#command-line-tools)
|
|
74
|
+
* [`pyard-import` Import Reference Data](#pyard-import-import-the-latest-ipd-imgthla-database)
|
|
75
|
+
* [`pyard-status` Show Statuses of Databases](#pyard-status-show-database-status)
|
|
76
|
+
* [`pyard` Redux](#pyard-redux-quickly)
|
|
77
|
+
* [`pyard-reduce-csv` Batch Mode Redux](#pyard-reduce-csv-batch-reduce-a-csv-file)
|
|
78
|
+
4. [`py-ard` REST Webservice](#py-ard-rest-web-service)
|
|
79
|
+
5. [Docker Deployment](#docker-deployment-of-py-ard-rest-web-service)
|
|
80
|
+
|
|
81
|
+
## Installation
|
|
82
|
+
|
|
83
|
+
`py-ard` works with Python 3.9 and higher (Python 3.8-3.13 are supported, but 3.9+ is recommended).
|
|
84
|
+
|
|
85
|
+
### Install from PyPi
|
|
86
|
+
|
|
87
|
+
```shell
|
|
88
|
+
pip install py-ard
|
|
89
|
+
```
|
|
90
|
+
|
|
91
|
+
### Install With Homebrew
|
|
92
|
+
|
|
93
|
+
On macOS, `py-ard` can be installed using Homebrew package manager.
|
|
94
|
+
This is very handy for using the command line versions of the tool without having to create virtual environments.
|
|
95
|
+
|
|
96
|
+
First time, you'd need to tap the `nmdp-bioinformatics` tap.
|
|
97
|
+
|
|
98
|
+
```shell
|
|
99
|
+
brew tap nmdp-bioinformatics/tap
|
|
100
|
+
```
|
|
101
|
+
|
|
102
|
+
Install `py-ard`
|
|
103
|
+
|
|
104
|
+
```shell
|
|
105
|
+
brew install py-ard
|
|
106
|
+
```
|
|
107
|
+
|
|
108
|
+
Homebrew will notify you as new versions of `py-ard` are released.
|
|
109
|
+
|
|
110
|
+
### Install from source
|
|
111
|
+
|
|
112
|
+
Checkout the `py-ard` source code.
|
|
113
|
+
|
|
114
|
+
```shell
|
|
115
|
+
git clone https://github.com/nmdp-bioinformatics/py-ard.git
|
|
116
|
+
cd py-ard
|
|
117
|
+
```
|
|
118
|
+
|
|
119
|
+
`py-ard` uses [uv](https://docs.astral.sh/uv/) to manage the project environment
|
|
120
|
+
and dependencies. Create the environment and install all dependency groups and
|
|
121
|
+
extras with:
|
|
122
|
+
|
|
123
|
+
```shell
|
|
124
|
+
make install
|
|
125
|
+
```
|
|
126
|
+
|
|
127
|
+
which runs `uv sync --all-extras --group test --group dev` and sets up
|
|
128
|
+
pre-commit. Run commands inside the environment with `uv run`, e.g.
|
|
129
|
+
`uv run pytest`.
|
|
130
|
+
|
|
131
|
+
See [Our Contribution Guide](CONTRIBUTING.rst) for open source contribution to `py-ard`.
|
|
132
|
+
|
|
133
|
+
## Using `py-ard`
|
|
134
|
+
|
|
135
|
+
### Using `py-ard` from Python code
|
|
136
|
+
|
|
137
|
+
`py-ard` can be used in a program to reduce/expand HLA GL String representation. If `py-ard` discovers an invalid Allele,
|
|
138
|
+
it'll throw an Invalid Exception, not silently return an empty result.
|
|
139
|
+
|
|
140
|
+
#### Initialize `py-ard`
|
|
141
|
+
|
|
142
|
+
Import and initialize `pyard` package.
|
|
143
|
+
The default initialization is to use the latest version of IPD-IMGT/HLA database.
|
|
144
|
+
|
|
145
|
+
```python
|
|
146
|
+
import pyard
|
|
147
|
+
|
|
148
|
+
ard = pyard.init()
|
|
149
|
+
```
|
|
150
|
+
|
|
151
|
+
Initialize `py-ard` with a particular version of IPD/IMGT-HLA database.
|
|
152
|
+
|
|
153
|
+
```python
|
|
154
|
+
import pyard
|
|
155
|
+
|
|
156
|
+
ard = pyard.init('3510')
|
|
157
|
+
```
|
|
158
|
+
|
|
159
|
+
When processing a large numbers of typings, it's helpful to have a cache of previously calculated reductions to make
|
|
160
|
+
similar typings reduce faster. The cache size of pre-computed reductions can be changed from the default of 1,000 by
|
|
161
|
+
setting `cache_size` argument. This increases the memory footprint but will significantly increase the processing times
|
|
162
|
+
for large number of reductions.
|
|
163
|
+
|
|
164
|
+
```python
|
|
165
|
+
import pyard
|
|
166
|
+
|
|
167
|
+
max_cache_size = 1_000_000
|
|
168
|
+
ard = pyard.init('3510', cache_size=max_cache_size)
|
|
169
|
+
```
|
|
170
|
+
|
|
171
|
+
By default, the IPD-IMGT/HLA data is stored locally in `$TMPDIR/pyard-$USER/`. This temporary location may be removed when your computer restarts.
|
|
172
|
+
|
|
173
|
+
Alternatively, you can specify a different, more permanent directory for the cached data.
|
|
174
|
+
|
|
175
|
+
```python
|
|
176
|
+
import pyard
|
|
177
|
+
|
|
178
|
+
ard = pyard.init('3510', data_dir='~/.py-ard/')
|
|
179
|
+
# Creating ~/.py-ard/pyard-3510.sqlite3 as cache.
|
|
180
|
+
# Version: 3510
|
|
181
|
+
```
|
|
182
|
+
|
|
183
|
+
As MAC data changes frequently, you can choose to refresh the MAC code for current IPD/IMGT-HLA database version.
|
|
184
|
+
|
|
185
|
+
```python
|
|
186
|
+
ard.refresh_mac_codes()
|
|
187
|
+
```
|
|
188
|
+
|
|
189
|
+
|
|
190
|
+
You can check the current version of IPD-IMGT/HLA database.
|
|
191
|
+
|
|
192
|
+
```python
|
|
193
|
+
ard.get_db_version()
|
|
194
|
+
```
|
|
195
|
+
|
|
196
|
+
You can choose to skip loading MAC codes if not needed (improves initialization time) by specifying `load_mac=False` during initialization.
|
|
197
|
+
|
|
198
|
+
```python
|
|
199
|
+
import pyard
|
|
200
|
+
|
|
201
|
+
ard = pyard.init('3510', load_mac=False)
|
|
202
|
+
```
|
|
203
|
+
|
|
204
|
+
#### Configure Reduction Behavior
|
|
205
|
+
|
|
206
|
+
Customize reduction behavior by passing a `config` dictionary to `pyard.init()`.
|
|
207
|
+
|
|
208
|
+
```python
|
|
209
|
+
import pyard
|
|
210
|
+
|
|
211
|
+
config = {
|
|
212
|
+
'reduce_serology': True, # Reduce serology typings (default: True)
|
|
213
|
+
'reduce_v2': True, # Reduce V2 alleles (default: True)
|
|
214
|
+
'reduce_3field': True, # Reduce 3-field alleles (default: True)
|
|
215
|
+
'reduce_P': True, # Reduce P group alleles (default: True)
|
|
216
|
+
'reduce_XX': True, # Reduce XX codes (default: True)
|
|
217
|
+
'reduce_MAC': True, # Reduce MAC codes (default: True)
|
|
218
|
+
'reduce_shortnull': True, # Reduce short nulls (default: True)
|
|
219
|
+
'ping': True, # Use ping mode (default: True)
|
|
220
|
+
'verbose_log': False, # Enable verbose logging (default: False)
|
|
221
|
+
'ARS_as_lg': False, # Treat ARS as lg (default: False)
|
|
222
|
+
'strict': True, # Strict validation mode (default: True)
|
|
223
|
+
'ignore_allele_with_suffixes': () # Tuple of suffixes to ignore (default: ())
|
|
224
|
+
}
|
|
225
|
+
|
|
226
|
+
ard = pyard.init('3510', config=config)
|
|
227
|
+
```
|
|
228
|
+
|
|
229
|
+
### `.pyardrc` Configuration File
|
|
230
|
+
|
|
231
|
+
`py-ard` looks for a `.pyardrc` file in the current directory first, then in the home directory (`~/`).
|
|
232
|
+
When found, it is used as the default configuration for `pyard.init()`,
|
|
233
|
+
so you can call `pyard.init()` with no arguments and have your preferred settings applied automatically.
|
|
234
|
+
|
|
235
|
+
The file uses [TOML](https://toml.io) format.
|
|
236
|
+
Copy [`extras/sample.pyardrc`](extras/sample.pyardrc) as a starting template:
|
|
237
|
+
|
|
238
|
+
```shell
|
|
239
|
+
cp extras/sample.pyardrc ~/.pyardrc
|
|
240
|
+
```
|
|
241
|
+
|
|
242
|
+
The `[pyard]` table maps to the `pyard.init()` parameters:
|
|
243
|
+
|
|
244
|
+
```toml
|
|
245
|
+
[pyard]
|
|
246
|
+
imgt_version = "3640"
|
|
247
|
+
data_dir = "~/.py-ard/"
|
|
248
|
+
load_mac = true
|
|
249
|
+
cache_size = 1000
|
|
250
|
+
```
|
|
251
|
+
|
|
252
|
+
The `[pyard.config]` table maps to the `ARDConfig` reduction settings:
|
|
253
|
+
|
|
254
|
+
```toml
|
|
255
|
+
[pyard.config]
|
|
256
|
+
reduce_serology = true
|
|
257
|
+
reduce_v2 = true
|
|
258
|
+
reduce_3field = true
|
|
259
|
+
reduce_P = true
|
|
260
|
+
reduce_XX = true
|
|
261
|
+
reduce_MAC = true
|
|
262
|
+
reduce_shortnull = true
|
|
263
|
+
ping = true
|
|
264
|
+
verbose_log = false
|
|
265
|
+
ARS_as_lg = false
|
|
266
|
+
strict = true
|
|
267
|
+
ignore_allele_with_suffixes = []
|
|
268
|
+
```
|
|
269
|
+
|
|
270
|
+
Any argument passed explicitly to `pyard.init()` takes precedence over the `.pyardrc` values:
|
|
271
|
+
|
|
272
|
+
You can set the environment variable `PYARD_RC` to `no` to skip reading the `.pyardrc` config.
|
|
273
|
+
|
|
274
|
+
```shell
|
|
275
|
+
export PYARD_RC=no
|
|
276
|
+
pyard -g "A*02:01" -r hats
|
|
277
|
+
```
|
|
278
|
+
|
|
279
|
+
### Reduce Typings
|
|
280
|
+
|
|
281
|
+
**Note**: The `redux` method in ARD object handles both GL Strings and individual alleles.
|
|
282
|
+
|
|
283
|
+
Reduce a single locus HLA Typing by specifying the allele/MAC/XX code and the reduction method to `redux`.
|
|
284
|
+
|
|
285
|
+
```python
|
|
286
|
+
allele = "A*01:01:01"
|
|
287
|
+
|
|
288
|
+
ard.redux(allele, 'G')
|
|
289
|
+
# >>> 'A*01:01:01G'
|
|
290
|
+
|
|
291
|
+
ard.redux(allele, 'lg')
|
|
292
|
+
# >>> 'A*01:01g'
|
|
293
|
+
|
|
294
|
+
ard.redux(allele, 'lgx')
|
|
295
|
+
# >>> 'A*01:01'
|
|
296
|
+
```
|
|
297
|
+
|
|
298
|
+
Reduce an ambiguous GL String
|
|
299
|
+
|
|
300
|
+
```python
|
|
301
|
+
# Reduce GL String
|
|
302
|
+
#
|
|
303
|
+
ard.redux("A*01:01/A*01:01N+A*02:AB^B*07:02+B*07:AB", "G")
|
|
304
|
+
# 'B*07:02:01G+B*07:02:01G^A*01:01:01G+A*02:01:01G/A*02:02'
|
|
305
|
+
```
|
|
306
|
+
|
|
307
|
+
You can also reduce serology based typings.
|
|
308
|
+
|
|
309
|
+
```python
|
|
310
|
+
ard.redux('B14', 'lg')
|
|
311
|
+
# >>> 'B*14:01g/B*14:02g/B*14:03g/B*14:04g/B*14:05g/B*14:06g/B*14:08g/B*14:09g/B*14:10g/B*14:11g/B*14:12g/B*14:13g/B*14:14g/B*14:15g/B*14:16g/B*14:17g/B*14:18g/B*14:19g/B*14:20g/B*14:21g/B*14:22g/B*14:23g/B*14:24g/B*14:25g/B*14:26g/B*14:27g/B*14:28g/B*14:29g/B*14:30g/B*14:31g/B*14:32g/B*14:33g/B*14:34g/B*14:35g/B*14:36g/B*14:37g/B*14:38g/B*14:39g/B*14:40g/B*14:42g/B*14:43g/B*14:44g/B*14:45g/B*14:46g/B*14:47g/B*14:48g/B*14:49g/B*14:50g/B*14:51g/B*14:52g/B*14:53g/B*14:54g/B*14:55g/B*14:56g/B*14:57g/B*14:58g/B*14:59g/B*14:60g/B*14:62g/B*14:63g/B*14:65g/B*14:66g/B*14:68g/B*14:70Qg/B*14:71g/B*14:73g/B*14:74g/B*14:75g/B*14:77g/B*14:82g/B*14:83g/B*14:86g/B*14:87g/B*14:88g/B*14:90g/B*14:93g/B*14:94g/B*14:95g/B*14:96g/B*14:97g/B*14:99g/B*14:102g'
|
|
312
|
+
```
|
|
313
|
+
|
|
314
|
+
## Valid Reduction Types
|
|
315
|
+
|
|
316
|
+
| Reduction Type | Description |
|
|
317
|
+
|----------------|-----------------------------------------------------------|
|
|
318
|
+
| `G` | Reduce to G Group Level |
|
|
319
|
+
| `P` | Reduce to P Group Level |
|
|
320
|
+
| `lg` | Reduce to 2 field ARD level (append `g`) |
|
|
321
|
+
| `lgx` | Reduce to 2 field ARD level |
|
|
322
|
+
| `W` | Reduce/Expand to full field(4,3,2) WHO nomenclature level |
|
|
323
|
+
| `exon` | Reduce/Expand to 3 field level |
|
|
324
|
+
| `U2` | Reduce to 2 field unambiguous level |
|
|
325
|
+
| `S` | Reduce to Serological level |
|
|
326
|
+
| `1F` | Reduce to First Field level |
|
|
327
|
+
| `hats` | Reduce to Antigen Specificity using HATS strategy |
|
|
328
|
+
|
|
329
|
+
### Perform DRB1 blending with DRB3, DRB4 and DRB5
|
|
330
|
+
|
|
331
|
+
```python
|
|
332
|
+
import pyard
|
|
333
|
+
|
|
334
|
+
pyard.dr_blender(drb1='HLA-DRB1*03:01+DRB1*04:01', drb3='DRB3*01:01', drb4='DRB4*01:03')
|
|
335
|
+
# >>> 'DRB3*01:01+DRB4*01:03'
|
|
336
|
+
```
|
|
337
|
+
|
|
338
|
+
## MAC Codes
|
|
339
|
+
|
|
340
|
+
`py-ard` supports not only reducing to various types but helps in expanding and
|
|
341
|
+
looking up MAC representation. See [MAC Service UI](https://hml.nmdp.org/MacUI/) for detail.
|
|
342
|
+
|
|
343
|
+
### Expand MAC
|
|
344
|
+
|
|
345
|
+
You can also use `py-ard` to expand MAC codes. Use `expand_mac` method on `ard`.
|
|
346
|
+
|
|
347
|
+
```python
|
|
348
|
+
ard.expand_mac('HLA-A*01:BC')
|
|
349
|
+
# 'HLA-A*01:02/HLA-A*01:03'
|
|
350
|
+
```
|
|
351
|
+
|
|
352
|
+
### Lookup MAC
|
|
353
|
+
|
|
354
|
+
Find the corresponding MAC code for an allele list GL String.
|
|
355
|
+
|
|
356
|
+
```python
|
|
357
|
+
ard.lookup_mac('A*01:02/A*01:01/A*01:03')
|
|
358
|
+
# A*01:MN
|
|
359
|
+
```
|
|
360
|
+
|
|
361
|
+
### CWD (Version 2) Reduction
|
|
362
|
+
|
|
363
|
+
Reduce a MAC code or an allele list GL String to CWD reduced list.
|
|
364
|
+
|
|
365
|
+
```python
|
|
366
|
+
ard.cwd_redux("B*15:01:01/B*15:01:03/B*15:04/B*15:07/B*15:26N/B*15:27")
|
|
367
|
+
# => B*15:01/B*15:07
|
|
368
|
+
```
|
|
369
|
+
|
|
370
|
+
The above 2 methods can be chained to get back a MAC code that has a CWD reduced version.
|
|
371
|
+
|
|
372
|
+
```python
|
|
373
|
+
ard.lookup_mac(ard.cwd_redux("B*15:01:01/B*15:01:03/B*15:04/B*15:07/B*15:26N/B*15:27"))
|
|
374
|
+
# 'B*15:AH'
|
|
375
|
+
```
|
|
376
|
+
|
|
377
|
+
### HATS reduction mode
|
|
378
|
+
|
|
379
|
+
Reduce to Antigen Specificity using HATS strategy
|
|
380
|
+
|
|
381
|
+
```python
|
|
382
|
+
ard.redux("B*44:450", "hats")
|
|
383
|
+
# => '4402'
|
|
384
|
+
```
|
|
385
|
+
|
|
386
|
+
### Additional Methods
|
|
387
|
+
|
|
388
|
+
Validate a GL String:
|
|
389
|
+
|
|
390
|
+
```python
|
|
391
|
+
ard.validate('A*01:01+A*02:01^B*07:02+B*08:01')
|
|
392
|
+
# Returns True if valid, raises exception if invalid
|
|
393
|
+
```
|
|
394
|
+
|
|
395
|
+
Expand XX codes:
|
|
396
|
+
|
|
397
|
+
```python
|
|
398
|
+
ard.expand_xx('A*01:XX')
|
|
399
|
+
# Returns all alleles matching the XX code
|
|
400
|
+
```
|
|
401
|
+
|
|
402
|
+
Find similar alleles:
|
|
403
|
+
|
|
404
|
+
```python
|
|
405
|
+
ard.similar_alleles('A*01:AB')
|
|
406
|
+
# Returns list of similar allele names
|
|
407
|
+
```
|
|
408
|
+
|
|
409
|
+
Check allele types:
|
|
410
|
+
|
|
411
|
+
```python
|
|
412
|
+
ard.is_mac('A*01:AB') # Check if MAC code
|
|
413
|
+
ard.is_serology('A1') # Check if serology
|
|
414
|
+
ard.is_v2('A*0101') # Check if V2 allele
|
|
415
|
+
ard.is_XX('A*01:XX') # Check if XX code
|
|
416
|
+
ard.is_shortnull('A*01:01N') # Check if short null
|
|
417
|
+
ard.is_null('A*01:01N') # Check if null allele
|
|
418
|
+
```
|
|
419
|
+
|
|
420
|
+
Validate an allele:
|
|
421
|
+
|
|
422
|
+
```python
|
|
423
|
+
ard.is_valid_allele('A*01:01:01') # True - valid 3-field allele
|
|
424
|
+
ard.is_valid_allele('A*01:01') # True - valid 2-field allele
|
|
425
|
+
ard.is_valid_allele('A*01:01:01G') # True - valid G group allele
|
|
426
|
+
ard.is_valid_allele('A*01:01P') # True - valid P group allele
|
|
427
|
+
ard.is_valid_allele('A*01:01g') # True - valid lg allele (non-strict mode)
|
|
428
|
+
ard.is_valid_allele('A*99:99') # False - allele not in database
|
|
429
|
+
```
|
|
430
|
+
|
|
431
|
+
`is_valid_allele` checks whether an allele exists in the IPD-IMGT/HLA database. It handles G group (suffix `G`), P group (suffix `P`), and lg (suffix `g`) allele designations. In strict mode, G and P group alleles are validated against their respective group mappings. For alleles with more than 2 fields, it falls back to checking the 2-field version if the full allele is not found.
|
|
432
|
+
|
|
433
|
+
Find serology relationships:
|
|
434
|
+
|
|
435
|
+
```python
|
|
436
|
+
ard.find_broad_splits('A10') # Find broad/split relationships
|
|
437
|
+
ard.find_associated_antigen('Bw4') # Find associated antigens
|
|
438
|
+
```
|
|
439
|
+
|
|
440
|
+
Convert V2 to V3:
|
|
441
|
+
|
|
442
|
+
```python
|
|
443
|
+
ard.v2_to_v3('A*0101') # Convert V2 allele to V3 format
|
|
444
|
+
```
|
|
445
|
+
|
|
446
|
+
### Using `py-ard` from R code
|
|
447
|
+
|
|
448
|
+
`py-ard` works well from `R` as well. Please
|
|
449
|
+
see [Using py-ard from R language](https://github.com/nmdp-bioinformatics/py-ard/wiki/Using-pyard-library-from-R-language)
|
|
450
|
+
page for detailed walkthrough.
|
|
451
|
+
|
|
452
|
+
## Command Line Tools
|
|
453
|
+
|
|
454
|
+
Various command line interface (CLI) tools are available to use for managing local IPD-IMGT/HLA cache database, running
|
|
455
|
+
impromptu reduction queries and batch processing of CSV files.
|
|
456
|
+
|
|
457
|
+
For all tools, use `--imgt-version` and `--data-dir` to specify the IPD-IMGT/HLA database version and the directory
|
|
458
|
+
where the SQLite files are created.
|
|
459
|
+
|
|
460
|
+
### `pyard-import` Import the latest IPD-IMGT/HLA database
|
|
461
|
+
|
|
462
|
+
`pyard-import` helps with importing and reinstalling of prepared IPD-IMGT/HLA and MAC data.
|
|
463
|
+
|
|
464
|
+
Use `pyard-import -h` to see all the options available.
|
|
465
|
+
|
|
466
|
+
```shell
|
|
467
|
+
$ pyard-import -h
|
|
468
|
+
usage: pyard-import [-h] [--list] [-i IPD_VERSION] [-d DATA_DIR]
|
|
469
|
+
[--v2-to-v3-mapping V2_V3_MAPPING] [--refresh-mac]
|
|
470
|
+
[--re-install] [--skip-mac]
|
|
471
|
+
|
|
472
|
+
py-ard tool to generate reference SQLite database. Allows updating db with
|
|
473
|
+
custom V2 to V3 mappings. Displays the list of available IPD/IMGT-HLA database
|
|
474
|
+
versions.
|
|
475
|
+
|
|
476
|
+
options:
|
|
477
|
+
-h, --help show this help message and exit
|
|
478
|
+
--list Show Versions of available IPD/IMGT-HLA Databases
|
|
479
|
+
-i, --ipd-version IPD_VERSION
|
|
480
|
+
Import supplied IPD/IMGT-HLA DB Version
|
|
481
|
+
-d, --data-dir DATA_DIR
|
|
482
|
+
Data directory to store imported data
|
|
483
|
+
--v2-to-v3-mapping V2_V3_MAPPING
|
|
484
|
+
V2 to V3 mapping CSV file
|
|
485
|
+
--refresh-mac Only refresh MAC data
|
|
486
|
+
--re-install reinstall a fresh version of database
|
|
487
|
+
--skip-mac Skip creating MAC mapping
|
|
488
|
+
```
|
|
489
|
+
|
|
490
|
+
Run `pyard-import` without any option to download and prepare the latest version of IPD-IMGT/HLA and MAC data.
|
|
491
|
+
|
|
492
|
+
```shell
|
|
493
|
+
$ pyard-import
|
|
494
|
+
Created Latest py-ard database
|
|
495
|
+
```
|
|
496
|
+
|
|
497
|
+
#### Import particular version of IPD/IMGT-HLA database
|
|
498
|
+
|
|
499
|
+
```shell
|
|
500
|
+
$ pyard-import --db-version 3.29.0
|
|
501
|
+
Created py-ard version 3290 database
|
|
502
|
+
```
|
|
503
|
+
|
|
504
|
+
Import particular version of IPD/IMGT-HLA database and replace the v2 to v3 mapping
|
|
505
|
+
table from a CSV file.
|
|
506
|
+
|
|
507
|
+
```shell
|
|
508
|
+
$ pyard-import --imgt-version 3.29.0 --v2-to-v3-mapping map2to3.csv
|
|
509
|
+
Created py-ard version 3290 database
|
|
510
|
+
Updated v2_mapping table with 'map2to3.csv' mapping file.
|
|
511
|
+
```
|
|
512
|
+
|
|
513
|
+
#### Reinstall a particular IPD/IMGT-HLA database
|
|
514
|
+
|
|
515
|
+
```shell
|
|
516
|
+
pyard-import --imgt-version 3340 --re-install
|
|
517
|
+
```
|
|
518
|
+
|
|
519
|
+
#### Replace the Latest IPD/IMGT-HLA database with V2 mappings
|
|
520
|
+
|
|
521
|
+
```shell
|
|
522
|
+
$ pyard-import --v2-to-v3-mapping map2to3.csv
|
|
523
|
+
```
|
|
524
|
+
|
|
525
|
+
#### Refresh the MAC for the specified version
|
|
526
|
+
|
|
527
|
+
```shell
|
|
528
|
+
$ pyard-import --imgt-version 3450 --refresh-mac
|
|
529
|
+
```
|
|
530
|
+
|
|
531
|
+
#### Skip MAC loading
|
|
532
|
+
|
|
533
|
+
You can skip loading MAC if you don't need by using `--skip-mac`
|
|
534
|
+
|
|
535
|
+
```shell
|
|
536
|
+
$ pyard-import --imgt-version 3150 --skip-mac
|
|
537
|
+
```
|
|
538
|
+
|
|
539
|
+
### `pyard-status` Show database status
|
|
540
|
+
|
|
541
|
+
Show the statuses of all `py-ard` databases
|
|
542
|
+
|
|
543
|
+
`pyard-status` goes through all the available databases and checks all the tables that should be available. This is very
|
|
544
|
+
helpful to show all the databases, number of rows in each table, any missing tables and the stored IPD-IMGT/HLA version.
|
|
545
|
+
|
|
546
|
+
```shell
|
|
547
|
+
$ pyard-status
|
|
548
|
+
```
|
|
549
|
+
|
|
550
|
+
Use ` --data-dir` to specify an alternate directory for cached database files.
|
|
551
|
+
|
|
552
|
+
```shell
|
|
553
|
+
$ pyard-status --data-dir ~/.pyard/
|
|
554
|
+
=============================================
|
|
555
|
+
IPD/IMGT-HLA DB Version: Latest (3530)
|
|
556
|
+
There is a newer IPD/IMGT-HLA release than version 3530
|
|
557
|
+
Upgrade to latest version '3630' with 'pyard-import --re-install'
|
|
558
|
+
File: /Users/pbashyal-nmdp/.pyard/pyard-Latest.sqlite3
|
|
559
|
+
Size: 577.42MB
|
|
560
|
+
---------------------------------------------
|
|
561
|
+
|Table Name | Rows|
|
|
562
|
+
|-------------------------------------------|
|
|
563
|
+
|alleles | 39,977|
|
|
564
|
+
|cwd2 | 336|
|
|
565
|
+
|dup_g | 70|
|
|
566
|
+
|exon_group | 13,406|
|
|
567
|
+
|exp_alleles | 91|
|
|
568
|
+
|g_group | 14,736|
|
|
569
|
+
|lgx_group | 14,736|
|
|
570
|
+
|mac_codes | 1,138,229|
|
|
571
|
+
|p_group | 21,534|
|
|
572
|
+
|p_not_g | 1,709|
|
|
573
|
+
|serology_broad_split_mapping | 23|
|
|
574
|
+
|serology_mapping | 131|
|
|
575
|
+
|shortnulls | 176|
|
|
576
|
+
|v2_mapping | 11|
|
|
577
|
+
|who_alleles | 37,619|
|
|
578
|
+
|who_group | 36,576|
|
|
579
|
+
|xx_codes | 2,019|
|
|
580
|
+
---------------------------------------------
|
|
581
|
+
```
|
|
582
|
+
|
|
583
|
+
### `pyard` Redux quickly
|
|
584
|
+
|
|
585
|
+
`pyard` command can be used for quick reductions from the command line. Use `--help` option to see all the available
|
|
586
|
+
options.
|
|
587
|
+
|
|
588
|
+
```shell
|
|
589
|
+
$ pyard --help
|
|
590
|
+
usage: pyard [-h] [-v] [-d DATA_DIR] [-i IPD_VERSION] [-g GL_STRING]
|
|
591
|
+
[-r {G,P,lg,lgx,W,exon,U2,S}] [--splits SPLITS] [--validate]
|
|
592
|
+
[--cwd CWD] [--expand-mac EXPAND_MAC] [--lookup-mac LOOKUP_MAC]
|
|
593
|
+
[--expand-xx EXPAND_XX] [--expand EXPAND]
|
|
594
|
+
[--similar SIMILAR_ALLELE] [--non-strict] [--verbose]
|
|
595
|
+
|
|
596
|
+
py-ard tool to redux GL String
|
|
597
|
+
|
|
598
|
+
options:
|
|
599
|
+
-h, --help show this help message and exit
|
|
600
|
+
-v, --version IPD-IMGT/HLA DB Version number
|
|
601
|
+
-d, --data-dir DATA_DIR
|
|
602
|
+
Data directory to store imported data
|
|
603
|
+
-i, --ipd-version IPD_VERSION
|
|
604
|
+
IPD-IMGT/HLA db to use for redux
|
|
605
|
+
-g, --gl GL_STRING GL String to reduce
|
|
606
|
+
-r, --redux-type {G,P,lg,lgx,W,exon,U2,S}
|
|
607
|
+
Reduction Method
|
|
608
|
+
--splits SPLITS Find Broad and Splits
|
|
609
|
+
--validate Validate the provided GL String
|
|
610
|
+
--cwd CWD Perform CWD redux
|
|
611
|
+
--expand-mac EXPAND_MAC
|
|
612
|
+
Expand MAC to Allele List
|
|
613
|
+
--lookup-mac LOOKUP_MAC
|
|
614
|
+
Lookup MAC for an Allele List
|
|
615
|
+
--expand-xx EXPAND_XX
|
|
616
|
+
Expand XX code to Allele List
|
|
617
|
+
--expand EXPAND Expand MAC or XX code to Allele List
|
|
618
|
+
--similar SIMILAR_ALLELE
|
|
619
|
+
Find Similar Alleles with given prefix
|
|
620
|
+
--non-strict Use non-strict mode
|
|
621
|
+
--verbose Use verbose mode
|
|
622
|
+
```
|
|
623
|
+
|
|
624
|
+
Reduce from command line by specifying any typing with `-g` or `--gl` option and the reduction method with `-r`
|
|
625
|
+
or `--redux-type` option.
|
|
626
|
+
|
|
627
|
+
```shell
|
|
628
|
+
$ pyard -g 'A*01:AB' -r lgx
|
|
629
|
+
A*01:01/A*01:02
|
|
630
|
+
|
|
631
|
+
$ pyard --gl 'DRB1*08:XX' -r G
|
|
632
|
+
DRB1*08:01:01G/DRB1*08:02:01G/DRB1*08:03:02G/DRB1*08:04:01G/DRB1*08:05/ ...
|
|
633
|
+
|
|
634
|
+
$ pyard -i 3290 --gl 'A1' -r lgx # For a particular version of DB
|
|
635
|
+
A*01:01/A*01:02/A*01:03/A*01:06/A*01:07/A*01:08/A*01:09/A*01:10/A*01:12/ ...
|
|
636
|
+
|
|
637
|
+
$ pyard -g "B*44:450" -r hats
|
|
638
|
+
4402
|
|
639
|
+
```
|
|
640
|
+
|
|
641
|
+
If the `-r` option is left out, `pyard` will print out the result of all reduction methods.
|
|
642
|
+
|
|
643
|
+
```shell
|
|
644
|
+
$ pyard -g 'A*01:01:01:01'
|
|
645
|
+
Reduction Method: G
|
|
646
|
+
-------------------
|
|
647
|
+
A*01:01:01G
|
|
648
|
+
|
|
649
|
+
Reduction Method: P
|
|
650
|
+
-------------------
|
|
651
|
+
A*01:01P
|
|
652
|
+
|
|
653
|
+
Reduction Method: lg
|
|
654
|
+
--------------------
|
|
655
|
+
A*01:01g
|
|
656
|
+
|
|
657
|
+
Reduction Method: lgx
|
|
658
|
+
---------------------
|
|
659
|
+
A*01:01
|
|
660
|
+
|
|
661
|
+
Reduction Method: W
|
|
662
|
+
-------------------
|
|
663
|
+
A*01:01:01:01
|
|
664
|
+
|
|
665
|
+
Reduction Method: exon
|
|
666
|
+
----------------------
|
|
667
|
+
A*01:01:01
|
|
668
|
+
|
|
669
|
+
Reduction Method: U2
|
|
670
|
+
--------------------
|
|
671
|
+
A*01:01
|
|
672
|
+
```
|
|
673
|
+
|
|
674
|
+
`py-ard` knows about the broad/splits of serology and DNA, you can find by using `--splits` option to `pyard` command.
|
|
675
|
+
|
|
676
|
+
```shell
|
|
677
|
+
$ pyard --splits "A*10"
|
|
678
|
+
A*10 = A*25/A*26/A*34/A*66
|
|
679
|
+
|
|
680
|
+
$ pyard --splits B14
|
|
681
|
+
B14 = B64/B65
|
|
682
|
+
```
|
|
683
|
+
|
|
684
|
+
Validate a GL String:
|
|
685
|
+
|
|
686
|
+
```shell
|
|
687
|
+
$ pyard -g 'A*01:01+A*02:01' --validate
|
|
688
|
+
```
|
|
689
|
+
|
|
690
|
+
Perform CWD reduction:
|
|
691
|
+
|
|
692
|
+
```shell
|
|
693
|
+
$ pyard --cwd 'B*15:01:01/B*15:01:03/B*15:04'
|
|
694
|
+
B*15:01
|
|
695
|
+
```
|
|
696
|
+
|
|
697
|
+
Expand MAC or XX codes:
|
|
698
|
+
|
|
699
|
+
```shell
|
|
700
|
+
$ pyard --expand-mac 'A*01:AB'
|
|
701
|
+
A*01:01/A*01:02
|
|
702
|
+
|
|
703
|
+
$ pyard --expand-xx 'A*01:XX'
|
|
704
|
+
A*01:01/A*01:02/A*01:03/...
|
|
705
|
+
```
|
|
706
|
+
Expand MAC based on HATS assignment for expanded alleles
|
|
707
|
+
|
|
708
|
+
```shell
|
|
709
|
+
$ pyard --expand-mac-hats "A*24:ABWMU"
|
|
710
|
+
````
|
|
711
|
+
|
|
712
|
+
Lookup MAC code:
|
|
713
|
+
|
|
714
|
+
```shell
|
|
715
|
+
$ pyard --lookup-mac 'A*01:01/A*01:02'
|
|
716
|
+
A*01:AB
|
|
717
|
+
```
|
|
718
|
+
|
|
719
|
+
Find similar alleles:
|
|
720
|
+
|
|
721
|
+
```shell
|
|
722
|
+
$ pyard --similar 'A*01:AB'
|
|
723
|
+
A*01:AB
|
|
724
|
+
A*01:AC
|
|
725
|
+
```
|
|
726
|
+
|
|
727
|
+
### `pyard-reduce-csv` Batch Reduce a CSV file
|
|
728
|
+
|
|
729
|
+
`pyard-reduce-csv` can be used to batch process a CSV file with HLA typings. See [documentation](extras/README.md) for
|
|
730
|
+
detailed information about all the options.
|
|
731
|
+
|
|
732
|
+
Generate sample configuration and CSV files:
|
|
733
|
+
|
|
734
|
+
```shell
|
|
735
|
+
$ pyard-reduce-csv --generate-sample
|
|
736
|
+
Created reduce_conf.json
|
|
737
|
+
Created sample.csv
|
|
738
|
+
Created reduce_conf_glstring.json
|
|
739
|
+
Created sample_glstring.csv
|
|
740
|
+
```
|
|
741
|
+
|
|
742
|
+
Reduce a CSV file using a configuration:
|
|
743
|
+
|
|
744
|
+
```shell
|
|
745
|
+
$ pyard-reduce-csv -c reduce_conf.json
|
|
746
|
+
```
|
|
747
|
+
|
|
748
|
+
## `py-ard` REST Web Service
|
|
749
|
+
|
|
750
|
+
Run `py-ard` as a service so that it can be accessed as a REST service endpoint.
|
|
751
|
+
|
|
752
|
+
To start in debug mode, you can run the `app.py` script. The endpoint should then be available
|
|
753
|
+
at [localhost:8080](http://0.0.0.0:8080)
|
|
754
|
+
|
|
755
|
+
```shell
|
|
756
|
+
$ python3 app.py
|
|
757
|
+
py-ard version: 2.0.0
|
|
758
|
+
IMGT version: 3631
|
|
759
|
+
`ConnexionMiddleware.run` is optimized for development. For production, run using a dedicated ASGI server.
|
|
760
|
+
INFO: Started server process [5344]
|
|
761
|
+
INFO: Waiting for application startup.
|
|
762
|
+
INFO: Application startup complete.
|
|
763
|
+
INFO: Uvicorn running on http://127.0.0.1:8080 (Press CTRL+C to quit)
|
|
764
|
+
```
|
|
765
|
+
|
|
766
|
+
## Docker deployment of py-ard REST Web Service
|
|
767
|
+
|
|
768
|
+
For deploying to production, build a Docker image and use that image for deploying to a server.
|
|
769
|
+
|
|
770
|
+
Build the docker image:
|
|
771
|
+
|
|
772
|
+
```shell
|
|
773
|
+
make docker-build
|
|
774
|
+
```
|
|
775
|
+
|
|
776
|
+
builds a Docker image named `nmdpbioinformatics/pyard-service:2.0.0.linux-amd64`
|
|
777
|
+
|
|
778
|
+
Build the docker and run it with:
|
|
779
|
+
|
|
780
|
+
```shell
|
|
781
|
+
make docker
|
|
782
|
+
```
|
|
783
|
+
|
|
784
|
+
The endpoint should then be available at [localhost:8080](http://0.0.0.0:8080)
|