salmopredict 0.1.2__tar.gz → 0.3.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- salmopredict-0.3.0/LICENSE +138 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/MANIFEST.in +1 -0
- salmopredict-0.3.0/PKG-INFO +239 -0
- salmopredict-0.3.0/README.md +213 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/environment.yml +1 -5
- salmopredict-0.3.0/examples/README.md +83 -0
- salmopredict-0.3.0/examples/example_gene_frequencies.csv +3 -0
- salmopredict-0.3.0/examples/example_samples.csv +11 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/pyproject.toml +6 -3
- {salmopredict-0.1.2 → salmopredict-0.3.0}/salmopredict/__init__.py +1 -1
- {salmopredict-0.1.2 → salmopredict-0.3.0}/salmopredict/cli.py +15 -4
- {salmopredict-0.1.2 → salmopredict-0.3.0}/salmopredict/config.py +12 -0
- salmopredict-0.3.0/salmopredict/core/frequencies.py +106 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/salmopredict/gui/app.py +67 -14
- {salmopredict-0.1.2 → salmopredict-0.3.0}/salmopredict/pipeline.py +31 -5
- salmopredict-0.3.0/salmopredict.egg-info/PKG-INFO +239 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/salmopredict.egg-info/SOURCES.txt +6 -1
- {salmopredict-0.1.2 → salmopredict-0.3.0}/salmopredict.egg-info/requires.txt +2 -1
- salmopredict-0.3.0/tests/test_two_table_input.py +144 -0
- salmopredict-0.1.2/LICENSE +0 -202
- salmopredict-0.1.2/PKG-INFO +0 -148
- salmopredict-0.1.2/README.md +0 -123
- salmopredict-0.1.2/salmopredict.egg-info/PKG-INFO +0 -148
- {salmopredict-0.1.2 → salmopredict-0.3.0}/examples/example_features.csv +0 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/examples/example_meta.csv +0 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/examples/example_with_sample.csv +0 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/salmopredict/core/__init__.py +0 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/salmopredict/core/align.py +0 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/salmopredict/core/io_tables.py +0 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/salmopredict/core/modelinfo.py +0 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/salmopredict/core/predict.py +0 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/salmopredict/gui/__init__.py +0 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/salmopredict/gui/assets/cfsa_logo.png +0 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/salmopredict/gui/assets/salmopredict_icon.png +0 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/salmopredict/gui/assets/vphs_logo.png +0 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/salmopredict/models/.DS_Store +0 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/salmopredict/models/model_default/learner.pkl +0 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/salmopredict/models/model_default/metadata.json +0 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/salmopredict/models/model_default/models/NeuralNetFastAI_r100_BAG_L1/S1F1/model-internals.pkl +0 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/salmopredict/models/model_default/models/NeuralNetFastAI_r100_BAG_L1/S1F1/model.pkl +0 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/salmopredict/models/model_default/models/NeuralNetFastAI_r100_BAG_L1/S1F2/model-internals.pkl +0 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/salmopredict/models/model_default/models/NeuralNetFastAI_r100_BAG_L1/S1F2/model.pkl +0 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/salmopredict/models/model_default/models/NeuralNetFastAI_r100_BAG_L1/S1F3/model-internals.pkl +0 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/salmopredict/models/model_default/models/NeuralNetFastAI_r100_BAG_L1/S1F3/model.pkl +0 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/salmopredict/models/model_default/models/NeuralNetFastAI_r100_BAG_L1/S1F4/model-internals.pkl +0 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/salmopredict/models/model_default/models/NeuralNetFastAI_r100_BAG_L1/S1F4/model.pkl +0 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/salmopredict/models/model_default/models/NeuralNetFastAI_r100_BAG_L1/S1F5/model-internals.pkl +0 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/salmopredict/models/model_default/models/NeuralNetFastAI_r100_BAG_L1/S1F5/model.pkl +0 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/salmopredict/models/model_default/models/NeuralNetFastAI_r100_BAG_L1/model.pkl +0 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/salmopredict/models/model_default/models/NeuralNetFastAI_r134_BAG_L1/S1F1/model-internals.pkl +0 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/salmopredict/models/model_default/models/NeuralNetFastAI_r134_BAG_L1/S1F1/model.pkl +0 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/salmopredict/models/model_default/models/NeuralNetFastAI_r134_BAG_L1/S1F2/model-internals.pkl +0 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/salmopredict/models/model_default/models/NeuralNetFastAI_r134_BAG_L1/S1F2/model.pkl +0 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/salmopredict/models/model_default/models/NeuralNetFastAI_r134_BAG_L1/S1F3/model-internals.pkl +0 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/salmopredict/models/model_default/models/NeuralNetFastAI_r134_BAG_L1/S1F3/model.pkl +0 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/salmopredict/models/model_default/models/NeuralNetFastAI_r134_BAG_L1/S1F4/model-internals.pkl +0 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/salmopredict/models/model_default/models/NeuralNetFastAI_r134_BAG_L1/S1F4/model.pkl +0 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/salmopredict/models/model_default/models/NeuralNetFastAI_r134_BAG_L1/S1F5/model-internals.pkl +0 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/salmopredict/models/model_default/models/NeuralNetFastAI_r134_BAG_L1/S1F5/model.pkl +0 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/salmopredict/models/model_default/models/NeuralNetFastAI_r134_BAG_L1/model.pkl +0 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/salmopredict/models/model_default/models/NeuralNetFastAI_r156_BAG_L1/S1F1/model-internals.pkl +0 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/salmopredict/models/model_default/models/NeuralNetFastAI_r156_BAG_L1/S1F1/model.pkl +0 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/salmopredict/models/model_default/models/NeuralNetFastAI_r156_BAG_L1/S1F2/model-internals.pkl +0 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/salmopredict/models/model_default/models/NeuralNetFastAI_r156_BAG_L1/S1F2/model.pkl +0 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/salmopredict/models/model_default/models/NeuralNetFastAI_r156_BAG_L1/S1F3/model-internals.pkl +0 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/salmopredict/models/model_default/models/NeuralNetFastAI_r156_BAG_L1/S1F3/model.pkl +0 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/salmopredict/models/model_default/models/NeuralNetFastAI_r156_BAG_L1/S1F4/model-internals.pkl +0 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/salmopredict/models/model_default/models/NeuralNetFastAI_r156_BAG_L1/S1F4/model.pkl +0 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/salmopredict/models/model_default/models/NeuralNetFastAI_r156_BAG_L1/S1F5/model-internals.pkl +0 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/salmopredict/models/model_default/models/NeuralNetFastAI_r156_BAG_L1/S1F5/model.pkl +0 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/salmopredict/models/model_default/models/NeuralNetFastAI_r156_BAG_L1/model.pkl +0 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/salmopredict/models/model_default/models/NeuralNetTorch_r121_BAG_L1/S1F1/model.pkl +0 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/salmopredict/models/model_default/models/NeuralNetTorch_r121_BAG_L1/S1F2/model.pkl +0 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/salmopredict/models/model_default/models/NeuralNetTorch_r121_BAG_L1/S1F3/model.pkl +0 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/salmopredict/models/model_default/models/NeuralNetTorch_r121_BAG_L1/S1F4/model.pkl +0 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/salmopredict/models/model_default/models/NeuralNetTorch_r121_BAG_L1/S1F5/model.pkl +0 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/salmopredict/models/model_default/models/NeuralNetTorch_r121_BAG_L1/model.pkl +0 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/salmopredict/models/model_default/models/NeuralNetTorch_r1_BAG_L1/S1F1/model.pkl +0 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/salmopredict/models/model_default/models/NeuralNetTorch_r1_BAG_L1/S1F2/model.pkl +0 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/salmopredict/models/model_default/models/NeuralNetTorch_r1_BAG_L1/S1F3/model.pkl +0 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/salmopredict/models/model_default/models/NeuralNetTorch_r1_BAG_L1/S1F4/model.pkl +0 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/salmopredict/models/model_default/models/NeuralNetTorch_r1_BAG_L1/S1F5/model.pkl +0 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/salmopredict/models/model_default/models/NeuralNetTorch_r1_BAG_L1/model.pkl +0 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/salmopredict/models/model_default/models/NeuralNetTorch_r30_BAG_L1/S1F1/model.pkl +0 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/salmopredict/models/model_default/models/NeuralNetTorch_r30_BAG_L1/S1F2/model.pkl +0 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/salmopredict/models/model_default/models/NeuralNetTorch_r30_BAG_L1/S1F3/model.pkl +0 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/salmopredict/models/model_default/models/NeuralNetTorch_r30_BAG_L1/S1F4/model.pkl +0 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/salmopredict/models/model_default/models/NeuralNetTorch_r30_BAG_L1/S1F5/model.pkl +0 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/salmopredict/models/model_default/models/NeuralNetTorch_r30_BAG_L1/model.pkl +0 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/salmopredict/models/model_default/models/NeuralNetTorch_r79_BAG_L1/S1F1/model.pkl +0 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/salmopredict/models/model_default/models/NeuralNetTorch_r79_BAG_L1/S1F2/model.pkl +0 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/salmopredict/models/model_default/models/NeuralNetTorch_r79_BAG_L1/S1F3/model.pkl +0 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/salmopredict/models/model_default/models/NeuralNetTorch_r79_BAG_L1/S1F4/model.pkl +0 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/salmopredict/models/model_default/models/NeuralNetTorch_r79_BAG_L1/S1F5/model.pkl +0 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/salmopredict/models/model_default/models/NeuralNetTorch_r79_BAG_L1/model.pkl +0 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/salmopredict/models/model_default/models/WeightedEnsemble_L2/model.pkl +0 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/salmopredict/models/model_default/models/trainer.pkl +0 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/salmopredict/models/model_default/predictor.pkl +0 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/salmopredict/models/model_default/version.txt +0 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/salmopredict.egg-info/dependency_links.txt +0 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/salmopredict.egg-info/entry_points.txt +0 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/salmopredict.egg-info/top_level.txt +0 -0
- {salmopredict-0.1.2 → salmopredict-0.3.0}/setup.cfg +0 -0
|
@@ -0,0 +1,138 @@
|
|
|
1
|
+
# PolyForm Noncommercial License 1.0.0
|
|
2
|
+
|
|
3
|
+
<https://polyformproject.org/licenses/noncommercial/1.0.0>
|
|
4
|
+
|
|
5
|
+
Required Notice: Copyright 2026 State Key Laboratory of Veterinary Public Health
|
|
6
|
+
and Safety, China Agricultural University
|
|
7
|
+
|
|
8
|
+
## Acceptance
|
|
9
|
+
|
|
10
|
+
In order to get any license under these terms, you must agree
|
|
11
|
+
to them as both strict obligations and conditions to all
|
|
12
|
+
your licenses.
|
|
13
|
+
|
|
14
|
+
## Copyright License
|
|
15
|
+
|
|
16
|
+
The licensor grants you a copyright license for the
|
|
17
|
+
software to do everything you might do with the software
|
|
18
|
+
that would otherwise infringe the licensor's copyright
|
|
19
|
+
in it for any permitted purpose. However, you may
|
|
20
|
+
only distribute the software according to [Distribution
|
|
21
|
+
License](#distribution-license) and make changes or new works
|
|
22
|
+
based on the software according to [Changes and New Works
|
|
23
|
+
License](#changes-and-new-works-license).
|
|
24
|
+
|
|
25
|
+
## Distribution License
|
|
26
|
+
|
|
27
|
+
The licensor grants you an additional copyright license
|
|
28
|
+
to distribute copies of the software. Your license
|
|
29
|
+
to distribute covers distributing the software with
|
|
30
|
+
changes and new works permitted by [Changes and New Works
|
|
31
|
+
License](#changes-and-new-works-license).
|
|
32
|
+
|
|
33
|
+
## Notices
|
|
34
|
+
|
|
35
|
+
You must ensure that anyone who gets a copy of any part of
|
|
36
|
+
the software from you also gets a copy of these terms or the
|
|
37
|
+
URL for them above, as well as copies of any plain-text lines
|
|
38
|
+
beginning with `Required Notice:` that the licensor provided
|
|
39
|
+
with the software. For example, if the licensor provided the
|
|
40
|
+
following notice in the software:
|
|
41
|
+
|
|
42
|
+
Required Notice: Copyright Yoyodyne, Inc. (http://example.com)
|
|
43
|
+
|
|
44
|
+
You must ensure that any plain-text lines beginning with
|
|
45
|
+
`Required Notice:` appear.
|
|
46
|
+
|
|
47
|
+
## Changes and New Works License
|
|
48
|
+
|
|
49
|
+
The licensor grants you an additional copyright license to
|
|
50
|
+
make changes and new works based on the software for any
|
|
51
|
+
permitted purpose.
|
|
52
|
+
|
|
53
|
+
## Patent License
|
|
54
|
+
|
|
55
|
+
The licensor grants you a patent license for the software that
|
|
56
|
+
covers patent claims the licensor can license, or becomes able
|
|
57
|
+
to license, that you would infringe by using the software.
|
|
58
|
+
|
|
59
|
+
## Noncommercial Purposes
|
|
60
|
+
|
|
61
|
+
Any noncommercial purpose is a permitted purpose.
|
|
62
|
+
|
|
63
|
+
## Personal Uses
|
|
64
|
+
|
|
65
|
+
Personal use for research, experiment, and testing for
|
|
66
|
+
the benefit of public knowledge, personal study, private
|
|
67
|
+
entertainment, hobby projects, amateur pursuits, or religious
|
|
68
|
+
observance, without any anticipated commercial application,
|
|
69
|
+
is use for a permitted purpose.
|
|
70
|
+
|
|
71
|
+
## Noncommercial Organizations
|
|
72
|
+
|
|
73
|
+
Use by any charitable organization, educational institution,
|
|
74
|
+
public research organization, public safety or health
|
|
75
|
+
organization, environmental protection organization,
|
|
76
|
+
or government institution is use for a permitted purpose
|
|
77
|
+
regardless of the source of funding or obligations resulting
|
|
78
|
+
from the funding.
|
|
79
|
+
|
|
80
|
+
## Fair Use
|
|
81
|
+
|
|
82
|
+
You may have "fair use" rights for the software under the
|
|
83
|
+
law. These terms do not limit them.
|
|
84
|
+
|
|
85
|
+
## No Other Rights
|
|
86
|
+
|
|
87
|
+
These terms do not allow you to sublicense or transfer any of
|
|
88
|
+
your licenses to anyone else, or prevent the licensor from
|
|
89
|
+
granting licenses to anyone else. These terms do not imply
|
|
90
|
+
any other licenses.
|
|
91
|
+
|
|
92
|
+
## Patent Defense
|
|
93
|
+
|
|
94
|
+
If you make any written claim that the software infringes or
|
|
95
|
+
contributes to infringement of any patent, your patent license
|
|
96
|
+
for the software granted under these terms ends immediately. If
|
|
97
|
+
your company makes such a claim, your patent license ends
|
|
98
|
+
immediately for work on behalf of your company.
|
|
99
|
+
|
|
100
|
+
## Violations
|
|
101
|
+
|
|
102
|
+
The first time you are notified in writing that you have
|
|
103
|
+
violated any of these terms, or done anything with the software
|
|
104
|
+
not covered by your licenses, your licenses can nonetheless
|
|
105
|
+
continue if you come into full compliance with these terms,
|
|
106
|
+
and take practical steps to correct past violations, within
|
|
107
|
+
32 days of receiving notice. Otherwise, all your licenses
|
|
108
|
+
end immediately.
|
|
109
|
+
|
|
110
|
+
## No Liability
|
|
111
|
+
|
|
112
|
+
***As far as the law allows, the software comes as is, without
|
|
113
|
+
any warranty or condition, and the licensor will not be liable
|
|
114
|
+
to you for any damages arising out of these terms or the use
|
|
115
|
+
or nature of the software, under any kind of legal claim.***
|
|
116
|
+
|
|
117
|
+
## Definitions
|
|
118
|
+
|
|
119
|
+
The **licensor** is the individual or entity offering these
|
|
120
|
+
terms, and the **software** is the software the licensor makes
|
|
121
|
+
available under these terms.
|
|
122
|
+
|
|
123
|
+
**You** refers to the individual or entity agreeing to these
|
|
124
|
+
terms.
|
|
125
|
+
|
|
126
|
+
**Your company** is any legal entity, sole proprietorship,
|
|
127
|
+
or other kind of organization that you work for, plus all
|
|
128
|
+
organizations that have control over, are under the control of,
|
|
129
|
+
or are under common control with that organization. **Control**
|
|
130
|
+
means ownership of substantially all the assets of an entity,
|
|
131
|
+
or the power to direct its management and policies by vote,
|
|
132
|
+
contract, or otherwise. Control can be direct or indirect.
|
|
133
|
+
|
|
134
|
+
**Your licenses** are all the licenses granted to you for the
|
|
135
|
+
software under these terms.
|
|
136
|
+
|
|
137
|
+
**Use** means anything you do with the software requiring one
|
|
138
|
+
of your licenses.
|
|
@@ -0,0 +1,239 @@
|
|
|
1
|
+
Metadata-Version: 2.4
|
|
2
|
+
Name: salmopredict
|
|
3
|
+
Version: 0.3.0
|
|
4
|
+
Summary: AutoGluon-based Incidence predictor for Salmonella virulence-factor gene-frequency features
|
|
5
|
+
Author-email: Dongyan Shao <563608176@qq.com>
|
|
6
|
+
License: PolyForm-Noncommercial-1.0.0
|
|
7
|
+
Project-URL: Homepage, https://github.com/shaodongyan/SalmoPredict
|
|
8
|
+
Project-URL: Repository, https://github.com/shaodongyan/SalmoPredict
|
|
9
|
+
Project-URL: Issues, https://github.com/shaodongyan/SalmoPredict/issues
|
|
10
|
+
Keywords: salmonella,incidence,autogluon,prediction,bioinformatics
|
|
11
|
+
Classifier: License :: Other/Proprietary License
|
|
12
|
+
Classifier: Programming Language :: Python :: 3
|
|
13
|
+
Classifier: Topic :: Scientific/Engineering :: Bio-Informatics
|
|
14
|
+
Requires-Python: <3.11,>=3.10
|
|
15
|
+
Description-Content-Type: text/markdown
|
|
16
|
+
License-File: LICENSE
|
|
17
|
+
Requires-Dist: autogluon.tabular[fastai]==1.1.1
|
|
18
|
+
Requires-Dist: setuptools<81
|
|
19
|
+
Requires-Dist: pandas>=2.0
|
|
20
|
+
Requires-Dist: openpyxl>=3.0
|
|
21
|
+
Requires-Dist: rich-argparse>=1.4
|
|
22
|
+
Requires-Dist: streamlit>=1.30
|
|
23
|
+
Provides-Extra: gui
|
|
24
|
+
Requires-Dist: streamlit>=1.30; extra == "gui"
|
|
25
|
+
Dynamic: license-file
|
|
26
|
+
|
|
27
|
+
<p align="center">
|
|
28
|
+
<img src="https://raw.githubusercontent.com/shaodongyan/SalmoPredict/main/salmopredict/gui/assets/salmopredict_icon.png" alt="salmopredict" width="200">
|
|
29
|
+
</p>
|
|
30
|
+
|
|
31
|
+
# salmopredict
|
|
32
|
+
|
|
33
|
+
AutoGluon-based **Incidence** predictor for *Salmonella* virulence-factor
|
|
34
|
+
gene-frequency features, with a command-line interface and a Streamlit GUI.
|
|
35
|
+
|
|
36
|
+
Given a feature table (rows = samples, columns = virulence-factor genes),
|
|
37
|
+
salmopredict aligns the columns to the features a pre-trained AutoGluon
|
|
38
|
+
`TabularPredictor` expects, runs the `WeightedEnsemble_L2` model, and writes a
|
|
39
|
+
single prediction file. It reproduces the alignment used by the original
|
|
40
|
+
`predict_autogluon.py`: column names are normalised R-`make.names`-style
|
|
41
|
+
(`/` and `-` become `.`), genes the model expects but the input lacks are filled
|
|
42
|
+
with `0` (a missing gene means frequency 0), and extra input columns are ignored.
|
|
43
|
+
|
|
44
|
+
**Feature CSV input/output** — the output columns depend on whether the input
|
|
45
|
+
has a `Sample` column:
|
|
46
|
+
|
|
47
|
+
| Input | Output columns |
|
|
48
|
+
|-------|----------------|
|
|
49
|
+
| No `Sample` column (features only) | `Incidence(%)` |
|
|
50
|
+
| Has a `Sample` column | `Sample`, `Incidence(%)` |
|
|
51
|
+
| Has a `Sample` column **and** `--attach meta.csv` | `Sample`, `Incidence(%)`, + the metadata's other columns |
|
|
52
|
+
|
|
53
|
+
Metadata is joined on the `Sample` key (the metadata CSV must also have a
|
|
54
|
+
`Sample` column), so attaching metadata requires a `Sample` column in the input.
|
|
55
|
+
|
|
56
|
+
## Install
|
|
57
|
+
|
|
58
|
+
salmopredict runs on **Python 3.10** and loads its model with **AutoGluon
|
|
59
|
+
1.1.1** — both are hard requirements, because the model is pickled with that
|
|
60
|
+
exact stack.
|
|
61
|
+
|
|
62
|
+
**New installation from this source directory.** Run these commands from the
|
|
63
|
+
directory containing `pyproject.toml`:
|
|
64
|
+
|
|
65
|
+
```bash
|
|
66
|
+
conda create -n salmopredict python=3.10
|
|
67
|
+
conda activate salmopredict
|
|
68
|
+
python -m pip install .
|
|
69
|
+
salmopredict check
|
|
70
|
+
```
|
|
71
|
+
|
|
72
|
+
This installs AutoGluon 1.1.1 with its **Torch and FastAI backends**, compatible
|
|
73
|
+
`setuptools<81`, the Streamlit GUI, and the bundled prediction model. The
|
|
74
|
+
backends are required by the bundled ensemble; base `autogluon.tabular` alone
|
|
75
|
+
does not install them. AutoGluon 1.1.1 also needs `pkg_resources`, which newer
|
|
76
|
+
setuptools releases no longer provide.
|
|
77
|
+
|
|
78
|
+
Both interfaces are then available:
|
|
79
|
+
|
|
80
|
+
```bash
|
|
81
|
+
salmopredict run -i features.csv -o results/ # command line
|
|
82
|
+
salmopredict gui # browser GUI
|
|
83
|
+
```
|
|
84
|
+
|
|
85
|
+
**Conda environment (alternative, from this source directory).** Pins Python 3.10 and installs
|
|
86
|
+
AutoGluon via pip inside the env (conda-installed AutoGluon does not resolve
|
|
87
|
+
cleanly for this project):
|
|
88
|
+
|
|
89
|
+
```bash
|
|
90
|
+
conda env create -f environment.yml
|
|
91
|
+
conda activate salmopredict
|
|
92
|
+
salmopredict check
|
|
93
|
+
```
|
|
94
|
+
|
|
95
|
+
**Editable / development install (from a clone).**
|
|
96
|
+
|
|
97
|
+
```bash
|
|
98
|
+
python -m pip install -e . # installs the CLI and the Streamlit GUI
|
|
99
|
+
```
|
|
100
|
+
|
|
101
|
+
**Install the updated local wheel.** In a Python 3.10 environment:
|
|
102
|
+
|
|
103
|
+
```bash
|
|
104
|
+
python -m pip install --upgrade dist/salmopredict-0.3.0-py3-none-any.whl
|
|
105
|
+
salmopredict check
|
|
106
|
+
```
|
|
107
|
+
|
|
108
|
+
**Update an existing source installation.** Stop a running GUI with `Ctrl+C`,
|
|
109
|
+
then run from this source directory:
|
|
110
|
+
|
|
111
|
+
```bash
|
|
112
|
+
conda activate salmopredict
|
|
113
|
+
python -m pip install --upgrade -e .
|
|
114
|
+
salmopredict check
|
|
115
|
+
salmopredict gui
|
|
116
|
+
```
|
|
117
|
+
|
|
118
|
+
**PyPI installation / upgrade.** In a Python 3.10 environment:
|
|
119
|
+
|
|
120
|
+
```bash
|
|
121
|
+
python -m pip install --upgrade "salmopredict>=0.3.0"
|
|
122
|
+
salmopredict check
|
|
123
|
+
```
|
|
124
|
+
|
|
125
|
+
Version 0.3.0 includes two-table prediction and installs the required Torch,
|
|
126
|
+
FastAI and compatible setuptools dependencies automatically.
|
|
127
|
+
|
|
128
|
+
If the page opens but prediction reports `No module named 'pkg_resources'`,
|
|
129
|
+
`torch`, or `fastai`, use the update/repair command above and restart the GUI.
|
|
130
|
+
To verify actual prediction from a source checkout (use a new output folder):
|
|
131
|
+
|
|
132
|
+
```bash
|
|
133
|
+
salmopredict run -i examples/example_features.csv -o results_install_check/
|
|
134
|
+
```
|
|
135
|
+
|
|
136
|
+
## The model
|
|
137
|
+
|
|
138
|
+
The prediction model is **already bundled** with salmopredict — both in this
|
|
139
|
+
repository and inside the PyPI wheel — at `salmopredict/models/model_default`, a
|
|
140
|
+
30 MB deployment. salmopredict uses it automatically, so the tool works out of
|
|
141
|
+
the box with no extra download or build step.
|
|
142
|
+
|
|
143
|
+
Model resolution order is `--model`, then `$SALMOPREDICT_MODEL`, then the single
|
|
144
|
+
directory under the package `models/` folder; with nothing specified it uses the
|
|
145
|
+
bundled `model_default`. Pass `--model /path/to/other` to run a different
|
|
146
|
+
AutoGluon model.
|
|
147
|
+
|
|
148
|
+
## Usage
|
|
149
|
+
|
|
150
|
+
Ready-to-run inputs live in [`examples/`](examples/) (see its README):
|
|
151
|
+
`example_features.csv` (Type 1, no `Sample`), `example_with_sample.csv`
|
|
152
|
+
(Type 2, with `Sample`), and `example_meta.csv` (metadata to attach). Features
|
|
153
|
+
are `gene_frequency × log10(CFU dose)`, matching how the model was trained. Try
|
|
154
|
+
one immediately:
|
|
155
|
+
|
|
156
|
+
```bash
|
|
157
|
+
salmopredict run -i examples/example_features.csv -o results/
|
|
158
|
+
```
|
|
159
|
+
|
|
160
|
+
```bash
|
|
161
|
+
# Features only -> output has just Incidence(%)
|
|
162
|
+
salmopredict run -i features.csv -o results/ --model /path/to/model
|
|
163
|
+
|
|
164
|
+
# With a Sample column -> output has Sample, Incidence(%)
|
|
165
|
+
salmopredict run -i examples/example_with_sample.csv -o results/
|
|
166
|
+
|
|
167
|
+
# Attach metadata joined on the Sample key -> Sample, Incidence(%), + meta columns
|
|
168
|
+
salmopredict run -i examples/example_with_sample.csv -o results/ \
|
|
169
|
+
--attach examples/example_meta.csv
|
|
170
|
+
|
|
171
|
+
# Launch the GUI, or check the environment/model
|
|
172
|
+
salmopredict gui
|
|
173
|
+
salmopredict check --model /path/to/model
|
|
174
|
+
```
|
|
175
|
+
|
|
176
|
+
Each feature-input run writes one `pred_<input-stem>.csv` to the output directory; the
|
|
177
|
+
prediction column is `Incidence(%)`. Features filled with `0` (genes the model
|
|
178
|
+
expects but the input lacks) are always reported, and a prominent warning
|
|
179
|
+
appears when more than `--missing-warn-frac` (default 0.3) of the model's
|
|
180
|
+
features are missing.
|
|
181
|
+
|
|
182
|
+
## Predict from samples and gene frequencies
|
|
183
|
+
|
|
184
|
+
Supply two CSV files instead of calculating features yourself:
|
|
185
|
+
|
|
186
|
+
* **Samples**: `Sample,dose_cfu,serotype`. `dose_cfu` contains raw CFU, e.g.
|
|
187
|
+
`1000`, not `3`. Each Sample must be nonblank and unique.
|
|
188
|
+
* **Gene frequencies**: `Serotype` plus one column per gene, one row per
|
|
189
|
+
serotype, with numeric frequencies from 0 to 1. This accepts the layout of
|
|
190
|
+
`02_gene_frequencies.csv` directly.
|
|
191
|
+
|
|
192
|
+
```bash
|
|
193
|
+
salmopredict run \
|
|
194
|
+
--samples examples/example_samples.csv \
|
|
195
|
+
--gene-frequencies examples/example_gene_frequencies.csv \
|
|
196
|
+
-o results_two_tables/
|
|
197
|
+
```
|
|
198
|
+
|
|
199
|
+
For each sample, the program looks up its serotype and calculates
|
|
200
|
+
`gene_frequency × log10(dose_cfu)`, then predicts incidence. It writes:
|
|
201
|
+
|
|
202
|
+
* `features_example_samples.csv`: `Sample` and the calculated gene features.
|
|
203
|
+
* `pred_example_samples.csv`: `Sample,dose_cfu,serotype,Incidence(%)`.
|
|
204
|
+
|
|
205
|
+
Sample order and identifier strings (including leading zeros) are preserved.
|
|
206
|
+
Required header names are case-insensitive. Serotype values match exactly after
|
|
207
|
+
trimming outer spaces; synonyms and spelling differences are not guessed.
|
|
208
|
+
Missing serotypes stop the run and list affected samples. Duplicate sample IDs
|
|
209
|
+
or serotypes, blank required values, nonfinite/nonpositive doses, and frequencies
|
|
210
|
+
outside [0, 1] also stop the run. Additional sample columns are ignored. Missing
|
|
211
|
+
model genes use the existing fill-and-warning behavior.
|
|
212
|
+
|
|
213
|
+
`--samples` and `--gene-frequencies` must be used together and cannot be combined
|
|
214
|
+
with `-i` or `--attach`. Use `--force` to replace existing output files.
|
|
215
|
+
|
|
216
|
+
In the **GUI**, choose **Samples + gene frequencies** under **Input mode**,
|
|
217
|
+
upload both CSVs, inspect their previews, choose an output folder and click
|
|
218
|
+
**Run prediction**. The result table and both CSV download buttons appear after
|
|
219
|
+
success. **Feature CSV** selects the existing single-table workflow.
|
|
220
|
+
|
|
221
|
+
The two example inputs were reconstructed from the matching sample/dose metadata
|
|
222
|
+
and dose-weighted features. See [examples/README.md](examples/README.md) for their
|
|
223
|
+
provenance and rounding tolerance.
|
|
224
|
+
|
|
225
|
+
## License
|
|
226
|
+
|
|
227
|
+
Licensed under the [PolyForm Noncommercial License 1.0.0](LICENSE): free to use,
|
|
228
|
+
modify, and share for any **noncommercial** purpose — including research,
|
|
229
|
+
teaching, and personal use, and by academic, government, public-health, and
|
|
230
|
+
other nonprofit organizations — but **commercial use is not permitted**.
|
|
231
|
+
Developed at the State Key Laboratory of Veterinary Public Health and Safety,
|
|
232
|
+
China Agricultural University, in collaboration with the China National Center
|
|
233
|
+
for Food Safety Risk Assessment (CFSA).
|
|
234
|
+
|
|
235
|
+
<p align="center">
|
|
236
|
+
<img src="https://raw.githubusercontent.com/shaodongyan/SalmoPredict/main/salmopredict/gui/assets/vphs_logo.png" alt="State Key Laboratory of Veterinary Public Health and Safety" height="80">
|
|
237
|
+
|
|
238
|
+
<img src="https://raw.githubusercontent.com/shaodongyan/SalmoPredict/main/salmopredict/gui/assets/cfsa_logo.png" alt="China National Center for Food Safety Risk Assessment" height="80">
|
|
239
|
+
</p>
|
|
@@ -0,0 +1,213 @@
|
|
|
1
|
+
<p align="center">
|
|
2
|
+
<img src="https://raw.githubusercontent.com/shaodongyan/SalmoPredict/main/salmopredict/gui/assets/salmopredict_icon.png" alt="salmopredict" width="200">
|
|
3
|
+
</p>
|
|
4
|
+
|
|
5
|
+
# salmopredict
|
|
6
|
+
|
|
7
|
+
AutoGluon-based **Incidence** predictor for *Salmonella* virulence-factor
|
|
8
|
+
gene-frequency features, with a command-line interface and a Streamlit GUI.
|
|
9
|
+
|
|
10
|
+
Given a feature table (rows = samples, columns = virulence-factor genes),
|
|
11
|
+
salmopredict aligns the columns to the features a pre-trained AutoGluon
|
|
12
|
+
`TabularPredictor` expects, runs the `WeightedEnsemble_L2` model, and writes a
|
|
13
|
+
single prediction file. It reproduces the alignment used by the original
|
|
14
|
+
`predict_autogluon.py`: column names are normalised R-`make.names`-style
|
|
15
|
+
(`/` and `-` become `.`), genes the model expects but the input lacks are filled
|
|
16
|
+
with `0` (a missing gene means frequency 0), and extra input columns are ignored.
|
|
17
|
+
|
|
18
|
+
**Feature CSV input/output** — the output columns depend on whether the input
|
|
19
|
+
has a `Sample` column:
|
|
20
|
+
|
|
21
|
+
| Input | Output columns |
|
|
22
|
+
|-------|----------------|
|
|
23
|
+
| No `Sample` column (features only) | `Incidence(%)` |
|
|
24
|
+
| Has a `Sample` column | `Sample`, `Incidence(%)` |
|
|
25
|
+
| Has a `Sample` column **and** `--attach meta.csv` | `Sample`, `Incidence(%)`, + the metadata's other columns |
|
|
26
|
+
|
|
27
|
+
Metadata is joined on the `Sample` key (the metadata CSV must also have a
|
|
28
|
+
`Sample` column), so attaching metadata requires a `Sample` column in the input.
|
|
29
|
+
|
|
30
|
+
## Install
|
|
31
|
+
|
|
32
|
+
salmopredict runs on **Python 3.10** and loads its model with **AutoGluon
|
|
33
|
+
1.1.1** — both are hard requirements, because the model is pickled with that
|
|
34
|
+
exact stack.
|
|
35
|
+
|
|
36
|
+
**New installation from this source directory.** Run these commands from the
|
|
37
|
+
directory containing `pyproject.toml`:
|
|
38
|
+
|
|
39
|
+
```bash
|
|
40
|
+
conda create -n salmopredict python=3.10
|
|
41
|
+
conda activate salmopredict
|
|
42
|
+
python -m pip install .
|
|
43
|
+
salmopredict check
|
|
44
|
+
```
|
|
45
|
+
|
|
46
|
+
This installs AutoGluon 1.1.1 with its **Torch and FastAI backends**, compatible
|
|
47
|
+
`setuptools<81`, the Streamlit GUI, and the bundled prediction model. The
|
|
48
|
+
backends are required by the bundled ensemble; base `autogluon.tabular` alone
|
|
49
|
+
does not install them. AutoGluon 1.1.1 also needs `pkg_resources`, which newer
|
|
50
|
+
setuptools releases no longer provide.
|
|
51
|
+
|
|
52
|
+
Both interfaces are then available:
|
|
53
|
+
|
|
54
|
+
```bash
|
|
55
|
+
salmopredict run -i features.csv -o results/ # command line
|
|
56
|
+
salmopredict gui # browser GUI
|
|
57
|
+
```
|
|
58
|
+
|
|
59
|
+
**Conda environment (alternative, from this source directory).** Pins Python 3.10 and installs
|
|
60
|
+
AutoGluon via pip inside the env (conda-installed AutoGluon does not resolve
|
|
61
|
+
cleanly for this project):
|
|
62
|
+
|
|
63
|
+
```bash
|
|
64
|
+
conda env create -f environment.yml
|
|
65
|
+
conda activate salmopredict
|
|
66
|
+
salmopredict check
|
|
67
|
+
```
|
|
68
|
+
|
|
69
|
+
**Editable / development install (from a clone).**
|
|
70
|
+
|
|
71
|
+
```bash
|
|
72
|
+
python -m pip install -e . # installs the CLI and the Streamlit GUI
|
|
73
|
+
```
|
|
74
|
+
|
|
75
|
+
**Install the updated local wheel.** In a Python 3.10 environment:
|
|
76
|
+
|
|
77
|
+
```bash
|
|
78
|
+
python -m pip install --upgrade dist/salmopredict-0.3.0-py3-none-any.whl
|
|
79
|
+
salmopredict check
|
|
80
|
+
```
|
|
81
|
+
|
|
82
|
+
**Update an existing source installation.** Stop a running GUI with `Ctrl+C`,
|
|
83
|
+
then run from this source directory:
|
|
84
|
+
|
|
85
|
+
```bash
|
|
86
|
+
conda activate salmopredict
|
|
87
|
+
python -m pip install --upgrade -e .
|
|
88
|
+
salmopredict check
|
|
89
|
+
salmopredict gui
|
|
90
|
+
```
|
|
91
|
+
|
|
92
|
+
**PyPI installation / upgrade.** In a Python 3.10 environment:
|
|
93
|
+
|
|
94
|
+
```bash
|
|
95
|
+
python -m pip install --upgrade "salmopredict>=0.3.0"
|
|
96
|
+
salmopredict check
|
|
97
|
+
```
|
|
98
|
+
|
|
99
|
+
Version 0.3.0 includes two-table prediction and installs the required Torch,
|
|
100
|
+
FastAI and compatible setuptools dependencies automatically.
|
|
101
|
+
|
|
102
|
+
If the page opens but prediction reports `No module named 'pkg_resources'`,
|
|
103
|
+
`torch`, or `fastai`, use the update/repair command above and restart the GUI.
|
|
104
|
+
To verify actual prediction from a source checkout (use a new output folder):
|
|
105
|
+
|
|
106
|
+
```bash
|
|
107
|
+
salmopredict run -i examples/example_features.csv -o results_install_check/
|
|
108
|
+
```
|
|
109
|
+
|
|
110
|
+
## The model
|
|
111
|
+
|
|
112
|
+
The prediction model is **already bundled** with salmopredict — both in this
|
|
113
|
+
repository and inside the PyPI wheel — at `salmopredict/models/model_default`, a
|
|
114
|
+
30 MB deployment. salmopredict uses it automatically, so the tool works out of
|
|
115
|
+
the box with no extra download or build step.
|
|
116
|
+
|
|
117
|
+
Model resolution order is `--model`, then `$SALMOPREDICT_MODEL`, then the single
|
|
118
|
+
directory under the package `models/` folder; with nothing specified it uses the
|
|
119
|
+
bundled `model_default`. Pass `--model /path/to/other` to run a different
|
|
120
|
+
AutoGluon model.
|
|
121
|
+
|
|
122
|
+
## Usage
|
|
123
|
+
|
|
124
|
+
Ready-to-run inputs live in [`examples/`](examples/) (see its README):
|
|
125
|
+
`example_features.csv` (Type 1, no `Sample`), `example_with_sample.csv`
|
|
126
|
+
(Type 2, with `Sample`), and `example_meta.csv` (metadata to attach). Features
|
|
127
|
+
are `gene_frequency × log10(CFU dose)`, matching how the model was trained. Try
|
|
128
|
+
one immediately:
|
|
129
|
+
|
|
130
|
+
```bash
|
|
131
|
+
salmopredict run -i examples/example_features.csv -o results/
|
|
132
|
+
```
|
|
133
|
+
|
|
134
|
+
```bash
|
|
135
|
+
# Features only -> output has just Incidence(%)
|
|
136
|
+
salmopredict run -i features.csv -o results/ --model /path/to/model
|
|
137
|
+
|
|
138
|
+
# With a Sample column -> output has Sample, Incidence(%)
|
|
139
|
+
salmopredict run -i examples/example_with_sample.csv -o results/
|
|
140
|
+
|
|
141
|
+
# Attach metadata joined on the Sample key -> Sample, Incidence(%), + meta columns
|
|
142
|
+
salmopredict run -i examples/example_with_sample.csv -o results/ \
|
|
143
|
+
--attach examples/example_meta.csv
|
|
144
|
+
|
|
145
|
+
# Launch the GUI, or check the environment/model
|
|
146
|
+
salmopredict gui
|
|
147
|
+
salmopredict check --model /path/to/model
|
|
148
|
+
```
|
|
149
|
+
|
|
150
|
+
Each feature-input run writes one `pred_<input-stem>.csv` to the output directory; the
|
|
151
|
+
prediction column is `Incidence(%)`. Features filled with `0` (genes the model
|
|
152
|
+
expects but the input lacks) are always reported, and a prominent warning
|
|
153
|
+
appears when more than `--missing-warn-frac` (default 0.3) of the model's
|
|
154
|
+
features are missing.
|
|
155
|
+
|
|
156
|
+
## Predict from samples and gene frequencies
|
|
157
|
+
|
|
158
|
+
Supply two CSV files instead of calculating features yourself:
|
|
159
|
+
|
|
160
|
+
* **Samples**: `Sample,dose_cfu,serotype`. `dose_cfu` contains raw CFU, e.g.
|
|
161
|
+
`1000`, not `3`. Each Sample must be nonblank and unique.
|
|
162
|
+
* **Gene frequencies**: `Serotype` plus one column per gene, one row per
|
|
163
|
+
serotype, with numeric frequencies from 0 to 1. This accepts the layout of
|
|
164
|
+
`02_gene_frequencies.csv` directly.
|
|
165
|
+
|
|
166
|
+
```bash
|
|
167
|
+
salmopredict run \
|
|
168
|
+
--samples examples/example_samples.csv \
|
|
169
|
+
--gene-frequencies examples/example_gene_frequencies.csv \
|
|
170
|
+
-o results_two_tables/
|
|
171
|
+
```
|
|
172
|
+
|
|
173
|
+
For each sample, the program looks up its serotype and calculates
|
|
174
|
+
`gene_frequency × log10(dose_cfu)`, then predicts incidence. It writes:
|
|
175
|
+
|
|
176
|
+
* `features_example_samples.csv`: `Sample` and the calculated gene features.
|
|
177
|
+
* `pred_example_samples.csv`: `Sample,dose_cfu,serotype,Incidence(%)`.
|
|
178
|
+
|
|
179
|
+
Sample order and identifier strings (including leading zeros) are preserved.
|
|
180
|
+
Required header names are case-insensitive. Serotype values match exactly after
|
|
181
|
+
trimming outer spaces; synonyms and spelling differences are not guessed.
|
|
182
|
+
Missing serotypes stop the run and list affected samples. Duplicate sample IDs
|
|
183
|
+
or serotypes, blank required values, nonfinite/nonpositive doses, and frequencies
|
|
184
|
+
outside [0, 1] also stop the run. Additional sample columns are ignored. Missing
|
|
185
|
+
model genes use the existing fill-and-warning behavior.
|
|
186
|
+
|
|
187
|
+
`--samples` and `--gene-frequencies` must be used together and cannot be combined
|
|
188
|
+
with `-i` or `--attach`. Use `--force` to replace existing output files.
|
|
189
|
+
|
|
190
|
+
In the **GUI**, choose **Samples + gene frequencies** under **Input mode**,
|
|
191
|
+
upload both CSVs, inspect their previews, choose an output folder and click
|
|
192
|
+
**Run prediction**. The result table and both CSV download buttons appear after
|
|
193
|
+
success. **Feature CSV** selects the existing single-table workflow.
|
|
194
|
+
|
|
195
|
+
The two example inputs were reconstructed from the matching sample/dose metadata
|
|
196
|
+
and dose-weighted features. See [examples/README.md](examples/README.md) for their
|
|
197
|
+
provenance and rounding tolerance.
|
|
198
|
+
|
|
199
|
+
## License
|
|
200
|
+
|
|
201
|
+
Licensed under the [PolyForm Noncommercial License 1.0.0](LICENSE): free to use,
|
|
202
|
+
modify, and share for any **noncommercial** purpose — including research,
|
|
203
|
+
teaching, and personal use, and by academic, government, public-health, and
|
|
204
|
+
other nonprofit organizations — but **commercial use is not permitted**.
|
|
205
|
+
Developed at the State Key Laboratory of Veterinary Public Health and Safety,
|
|
206
|
+
China Agricultural University, in collaboration with the China National Center
|
|
207
|
+
for Food Safety Risk Assessment (CFSA).
|
|
208
|
+
|
|
209
|
+
<p align="center">
|
|
210
|
+
<img src="https://raw.githubusercontent.com/shaodongyan/SalmoPredict/main/salmopredict/gui/assets/vphs_logo.png" alt="State Key Laboratory of Veterinary Public Health and Safety" height="80">
|
|
211
|
+
|
|
212
|
+
<img src="https://raw.githubusercontent.com/shaodongyan/SalmoPredict/main/salmopredict/gui/assets/cfsa_logo.png" alt="China National Center for Food Safety Risk Assessment" height="80">
|
|
213
|
+
</p>
|
|
@@ -0,0 +1,83 @@
|
|
|
1
|
+
# Example inputs
|
|
2
|
+
|
|
3
|
+
Three ready-to-run files that cover the two input types and the metadata attach.
|
|
4
|
+
Feature values are built the way the model was trained:
|
|
5
|
+
|
|
6
|
+
```
|
|
7
|
+
feature = gene_frequency(serotype, gene) × log10(CFU dose)
|
|
8
|
+
```
|
|
9
|
+
|
|
10
|
+
so the same serotype at a higher CFU has larger feature values and a higher
|
|
11
|
+
predicted incidence. All files share the same 10 samples — two serotypes
|
|
12
|
+
(Enteritidis, Typhimurium) each at CFU = 500 / 1 000 / 2 000 / 10 000 / 100 000.
|
|
13
|
+
Gene column names keep the original biological form (`mig-5`, `spiC/ssaB`, …);
|
|
14
|
+
salmopredict normalises them to the model's names (`mig-5` → `mig.5`).
|
|
15
|
+
|
|
16
|
+
| File | Rows × cols | Type | Output when run |
|
|
17
|
+
|------|-------------|------|-----------------|
|
|
18
|
+
| `example_features.csv` | 10 × 123 | **Type 1** — features only, **no `Sample`** column (just the 123 genes the model uses) | `Incidence(%)` |
|
|
19
|
+
| `example_with_sample.csv` | 10 × 348 | **Type 2** — a `Sample` column + all 347 genes (the ~224 extra genes are ignored) | `Sample`, `Incidence(%)` |
|
|
20
|
+
| `example_meta.csv` | 10 × 5 | Metadata to **attach** — `Sample` + `serotype`, `dose_cfu`, `source`, `region` | joined onto a Type-2 run by the `Sample` key |
|
|
21
|
+
|
|
22
|
+
## Run them
|
|
23
|
+
|
|
24
|
+
```bash
|
|
25
|
+
# Type 1: features only -> a single Incidence(%) column
|
|
26
|
+
salmopredict run -i examples/example_features.csv -o results/
|
|
27
|
+
|
|
28
|
+
# Type 2: a Sample column -> Sample, Incidence(%)
|
|
29
|
+
salmopredict run -i examples/example_with_sample.csv -o results/
|
|
30
|
+
|
|
31
|
+
# Type 2 + attach: metadata joined on Sample -> Sample, Incidence(%), + meta columns
|
|
32
|
+
salmopredict run -i examples/example_with_sample.csv -o results/ \
|
|
33
|
+
--attach examples/example_meta.csv
|
|
34
|
+
```
|
|
35
|
+
|
|
36
|
+
The attached run shows the dose–response, since `example_meta.csv` carries the
|
|
37
|
+
`dose_cfu` alongside each `Sample`:
|
|
38
|
+
|
|
39
|
+
```
|
|
40
|
+
Sample,Incidence(%),serotype,dose_cfu,source,region
|
|
41
|
+
S001,12.39,Enteritidis,500,retail chicken,North
|
|
42
|
+
S002,14.15,Enteritidis,1000,retail pork,East
|
|
43
|
+
...
|
|
44
|
+
S005,37.43,Enteritidis,100000,retail pork,West
|
|
45
|
+
```
|
|
46
|
+
|
|
47
|
+
Rules for attaching metadata:
|
|
48
|
+
|
|
49
|
+
* the **input** must have a `Sample` column (Type 2), and so must the
|
|
50
|
+
**metadata** file — the two are joined on `Sample`;
|
|
51
|
+
* the metadata's `Sample` values should be unique (duplicates are rejected);
|
|
52
|
+
* every metadata column except `Sample` is appended to the output.
|
|
53
|
+
|
|
54
|
+
In the GUI (`salmopredict gui`), the **Attach metadata** box only appears once
|
|
55
|
+
the chosen input is detected to have a `Sample` column.
|
|
56
|
+
|
|
57
|
+
## Rebuilding these files
|
|
58
|
+
|
|
59
|
+
The values come from the per-serotype gene-frequency table used to train the
|
|
60
|
+
model (`results/02_gene_frequencies.csv` in the assembly project) multiplied by
|
|
61
|
+
`log10(dose)`, reproducing `multiply_CFU_geneFreq.R`. To change the serotypes or
|
|
62
|
+
CFU values, edit that grid and recompute `gene_frequency × log10(CFU)` for every
|
|
63
|
+
gene column — do not edit a dose value alone, or the features and the dose will
|
|
64
|
+
disagree.
|
|
65
|
+
|
|
66
|
+
## Two-table input examples
|
|
67
|
+
|
|
68
|
+
`example_samples.csv` contains 10 samples with raw `dose_cfu` and `serotype`.
|
|
69
|
+
`example_gene_frequencies.csv` has 2 serotypes and 123 model gene columns.
|
|
70
|
+
Both are also saved in the sibling `../test/` folder requested for verification.
|
|
71
|
+
|
|
72
|
+
These files were reconstructed from `../test/pred_example_with_sample.csv`
|
|
73
|
+
(sample IDs, serotypes and doses only) and `../test/example_features.csv`
|
|
74
|
+
(weighted features), pairing rows in their original order. The prediction column
|
|
75
|
+
is not used to derive frequencies. For each serotype, the 1000-CFU row gives
|
|
76
|
+
`frequency = feature / log10(1000) = feature / 3`. All five dose rows per serotype
|
|
77
|
+
were checked against that frequency. The maximum reconstructed feature error
|
|
78
|
+
is less than 5e-7, consistent with six-decimal rounding in the original features.
|
|
79
|
+
|
|
80
|
+
```bash
|
|
81
|
+
salmopredict run --samples examples/example_samples.csv \
|
|
82
|
+
--gene-frequencies examples/example_gene_frequencies.csv -o results_two_tables/
|
|
83
|
+
```
|