MetaPont 0.0.2__tar.gz → 0.0.3__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- metapont-0.0.3/PKG-INFO +210 -0
- metapont-0.0.3/README.md +195 -0
- {metapont-0.0.2 → metapont-0.0.3}/setup.cfg +2 -1
- {metapont-0.0.2 → metapont-0.0.3}/src/MetaPont/Extract_By_Function.py +28 -11
- metapont-0.0.3/src/MetaPont/Extract_By_Taxa.py +130 -0
- metapont-0.0.3/src/MetaPont/constants.py +2 -0
- metapont-0.0.3/src/MetaPont.egg-info/PKG-INFO +210 -0
- {metapont-0.0.2 → metapont-0.0.3}/src/MetaPont.egg-info/SOURCES.txt +1 -0
- {metapont-0.0.2 → metapont-0.0.3}/src/MetaPont.egg-info/entry_points.txt +1 -0
- metapont-0.0.2/PKG-INFO +0 -130
- metapont-0.0.2/README.md +0 -115
- metapont-0.0.2/src/MetaPont/constants.py +0 -2
- metapont-0.0.2/src/MetaPont.egg-info/PKG-INFO +0 -130
- {metapont-0.0.2 → metapont-0.0.3}/LICENSE +0 -0
- {metapont-0.0.2 → metapont-0.0.3}/pyproject.toml +0 -0
- {metapont-0.0.2 → metapont-0.0.3}/src/MetaPont/Function_By_Taxa.py +0 -0
- {metapont-0.0.2 → metapont-0.0.3}/src/MetaPont/__init__.py +0 -0
- {metapont-0.0.2 → metapont-0.0.3}/src/MetaPont/emapper_readmap_cds_combiner_v1.py +0 -0
- {metapont-0.0.2 → metapont-0.0.3}/src/MetaPont/kraken_emapper_readmap_combiner_v1.py +0 -0
- {metapont-0.0.2 → metapont-0.0.3}/src/MetaPont/lineage_matrix_collapse_v1.py +0 -0
- {metapont-0.0.2 → metapont-0.0.3}/src/MetaPont/lineage_matrix_collapse_v2.py +0 -0
- {metapont-0.0.2 → metapont-0.0.3}/src/MetaPont.egg-info/dependency_links.txt +0 -0
- {metapont-0.0.2 → metapont-0.0.3}/src/MetaPont.egg-info/top_level.txt +0 -0
metapont-0.0.3/PKG-INFO
ADDED
|
@@ -0,0 +1,210 @@
|
|
|
1
|
+
Metadata-Version: 2.1
|
|
2
|
+
Name: MetaPont
|
|
3
|
+
Version: 0.0.3
|
|
4
|
+
Summary: MetaPont - A tool to bridge the gap between the output of metagenomic tools and the analysis of the data
|
|
5
|
+
Home-page: https://github.com/TheHuwsLab/MetaPont
|
|
6
|
+
Author: Nicholas Dimonaco
|
|
7
|
+
Author-email: nicholas@dimonaco.co.uk
|
|
8
|
+
Project-URL: Bug Tracker, https://github.com/TheHuwsLab/MetaPont/issues
|
|
9
|
+
Classifier: Programming Language :: Python :: 3
|
|
10
|
+
Classifier: License :: OSI Approved :: GNU General Public License v3 (GPLv3)
|
|
11
|
+
Classifier: Operating System :: OS Independent
|
|
12
|
+
Requires-Python: >=3.6
|
|
13
|
+
Description-Content-Type: text/markdown
|
|
14
|
+
License-File: LICENSE
|
|
15
|
+
|
|
16
|
+
# MetaPont
|
|
17
|
+
**MetaPont** - A tool to bridge the gap between the output of metagenomic tools and the analysis of the data
|
|
18
|
+
|
|
19
|
+
## Features - These are the current aims of this project - Still under development
|
|
20
|
+
|
|
21
|
+
- **Targeted Functional Analysis:** Search for specific functional IDs (e.g., GO terms) within the `_Final_Contig.tsv` files provided by the HuwsLab Metagenome Workflow (https://github.com/TheHuwsLab/Metagenome_Workflow) .
|
|
22
|
+
- **Taxonomic Breakdown:** Extract genus-level taxonomy information and calculate their proportions in the dataset.
|
|
23
|
+
- **Batch Processing:** Analyse all `_Final_Contig.tsv` files in a specified directory.
|
|
24
|
+
- **Customisable Output:** Save results in a format suitable for downstream analysis.
|
|
25
|
+
|
|
26
|
+
---
|
|
27
|
+
|
|
28
|
+
## Installation
|
|
29
|
+
|
|
30
|
+
### Prerequisites
|
|
31
|
+
|
|
32
|
+
Ensure you have the following installed:
|
|
33
|
+
|
|
34
|
+
- Python ~3.6 or later
|
|
35
|
+
|
|
36
|
+
### Installation via pip
|
|
37
|
+
|
|
38
|
+
MetaPont is provided as a pip distribution.
|
|
39
|
+
|
|
40
|
+
```bash
|
|
41
|
+
pip install MetaPont
|
|
42
|
+
```
|
|
43
|
+
|
|
44
|
+
---
|
|
45
|
+
|
|
46
|
+
## Usage
|
|
47
|
+
|
|
48
|
+
---
|
|
49
|
+
### Extract-By-Function Command-line Arguments
|
|
50
|
+
```Extract-By-Function -h ```
|
|
51
|
+
```bash
|
|
52
|
+
usage: Extract_By_Function.py [-h] -d DIRECTORY -f FUNCTION_ID -o OUTPUT
|
|
53
|
+
[-m MIN_PROPORTION] [-top TOP_TAXA]
|
|
54
|
+
|
|
55
|
+
MetaPont v0.0.3: Extract-By-Function - Identify taxa contributing to a
|
|
56
|
+
specific function.
|
|
57
|
+
|
|
58
|
+
options:
|
|
59
|
+
-h, --help show this help message and exit
|
|
60
|
+
-d DIRECTORY, --directory DIRECTORY
|
|
61
|
+
Directory containing TSV files to analyse.
|
|
62
|
+
-f FUNCTION_ID, --function_id FUNCTION_ID
|
|
63
|
+
Specific function ID to search for (e.g.,
|
|
64
|
+
'GO:0016597').
|
|
65
|
+
-o OUTPUT, --output OUTPUT
|
|
66
|
+
Output file to save results.
|
|
67
|
+
-m MIN_PROPORTION, --min_proportion MIN_PROPORTION
|
|
68
|
+
Minimum proportion threshold for taxa to be included
|
|
69
|
+
in the output.
|
|
70
|
+
-top TOP_TAXA, --top_taxa TOP_TAXA
|
|
71
|
+
Top n taxa to be included in the output.
|
|
72
|
+
|
|
73
|
+
```
|
|
74
|
+
|
|
75
|
+
The `Extract-By-Function` tool provides several command-line options: \
|
|
76
|
+
Note: Either -m or -top is required.
|
|
77
|
+
|
|
78
|
+
| Option | Description | Required | Default |
|
|
79
|
+
|--------------------------|------------------------------------------------------------|----------|---------|
|
|
80
|
+
| `-d`, `--directory` | Directory containing `_Final_Contig.tsv` files to analyse. | Yes | None |
|
|
81
|
+
| `-f`, `--function_id` | Functional ID to search for (e.g., `GO:0016597`). | Yes | None |
|
|
82
|
+
| `-m`, `--min_proportion` | Minimum proportion needed for reporting. | Yes/No | None |
|
|
83
|
+
| `-top`, `--top_taxa` | Number of taxa to report. | Yes/No | None |
|
|
84
|
+
| `-o`, `--output` | Output file name to save results. | Yes | None |
|
|
85
|
+
|
|
86
|
+
### Example
|
|
87
|
+
|
|
88
|
+
To search for the functional ID `GO:0016597` in all `_Final_Contig.tsv` files within the `test_data/` directory:
|
|
89
|
+
|
|
90
|
+
```bash
|
|
91
|
+
Extract=By-Function -d .../test_data/Final_contig/ -f GO:0016597 -top 3 -o .../test_data/Final_Contig/Extract_By_Function_Out/results.tsv
|
|
92
|
+
```
|
|
93
|
+
|
|
94
|
+
---
|
|
95
|
+
|
|
96
|
+
## Output
|
|
97
|
+
|
|
98
|
+
The tool generates a tab-delimited output file with the following columns:
|
|
99
|
+
|
|
100
|
+
1. **Sample:** Name of the processed Sample.
|
|
101
|
+
2. **Taxa:** Genus-level taxonomic assignment extracted from the `Lineage` column.
|
|
102
|
+
3. **Reads Assigned (Function):** Number of reads assigned to contigs with the given functional ID.
|
|
103
|
+
3. **Proportion:** Proportion of reads assigned to contigs of stated Taxa with the given functional ID within the sample.
|
|
104
|
+
4. **Proportion (Total Reads):** Proportion of reads assigned to contigs of stated Taxa with the given functional ID within the total reads of the sample.
|
|
105
|
+
|
|
106
|
+
Example output:
|
|
107
|
+
|
|
108
|
+
```
|
|
109
|
+
Function ID: GO:0016597
|
|
110
|
+
Sample Taxa Reads Assigned (Function) Proportion (Function) Proportion (Total Reads)
|
|
111
|
+
PN0536_0001_S1_Final_Contig.tsv Lactobacillus 111963 0.602 0.004
|
|
112
|
+
PN0536_0003_S83_Final_Contig.tsv Lactobacillus 20072 0.457 0.001
|
|
113
|
+
PN0536_0002_S2_Final_Contig.tsv Acutalibacter 145222 0.795 0.005
|
|
114
|
+
PN0536_0004_S3_Final_Contig.tsv Lactobacillus 40076 0.404 0.002
|
|
115
|
+
```
|
|
116
|
+
|
|
117
|
+
---
|
|
118
|
+
|
|
119
|
+
|
|
120
|
+
|
|
121
|
+
### Workflow - unfinished
|
|
122
|
+
|
|
123
|
+
1. The script reads `_Final_Contig.tsv` files from the specified directory.
|
|
124
|
+
2. For each file, it searches for occurrences of the given functional ID within specific columns.
|
|
125
|
+
3. Matches are associated with genus-level taxonomic information extracted from the `Lineage` column.
|
|
126
|
+
4. Taxa proportions are calculated and saved to the output file.
|
|
127
|
+
|
|
128
|
+
---
|
|
129
|
+
## Extract-By-Taxa Command-line Arguments
|
|
130
|
+
```Extract-By-Taxa -h ```
|
|
131
|
+
```bash
|
|
132
|
+
usage: Extract_By_Taxa.py [-h] -d DIRECTORY -t TAXON -o OUTPUT -func
|
|
133
|
+
FUNCTIONAL_CLASSES [-top TOP_FUNCTIONS]
|
|
134
|
+
|
|
135
|
+
MetaPont: Extract Top Functions by Taxon
|
|
136
|
+
|
|
137
|
+
options:
|
|
138
|
+
-h, --help show this help message and exit
|
|
139
|
+
-d DIRECTORY, --directory DIRECTORY
|
|
140
|
+
Directory containing TSV files to analyse.
|
|
141
|
+
-t TAXON, --taxon TAXON
|
|
142
|
+
Target taxon to search for (e.g., 'g__Bacillus').
|
|
143
|
+
-o OUTPUT, --output OUTPUT
|
|
144
|
+
Output file to save results.
|
|
145
|
+
-func FUNCTIONAL_CLASSES, --functional_classes FUNCTIONAL_CLASSES
|
|
146
|
+
Which functional classes to report (e.g. GO,EC,KEGG
|
|
147
|
+
etc).
|
|
148
|
+
-top TOP_FUNCTIONS, --top_functions TOP_FUNCTIONS
|
|
149
|
+
Top n functions to include in the output for each
|
|
150
|
+
sample (default: 3).
|
|
151
|
+
|
|
152
|
+
```
|
|
153
|
+
|
|
154
|
+
The `Extract-By-Taxa` tool provides several command-line options:
|
|
155
|
+
|
|
156
|
+
|
|
157
|
+
| Option | Description | Required | Default |
|
|
158
|
+
|-----------------------------------|-------------------------------------------------------------|----------|---------|
|
|
159
|
+
| `-d`, `--directory` | Directory containing `_Fincal_Contig.tsv` files to analyse. | Yes | None |
|
|
160
|
+
| `-t`, `--taxon` | Taxa to search for (e.g., `g__Bacillus`). | Yes | None |
|
|
161
|
+
| `-func`, `--functional_classes` | Functional classes to report (e.g. GO,EC,KEGG etc). | Yes | None |
|
|
162
|
+
| `-top`, `--top_taxa` | Number of functions to report (default 3). | No | None |
|
|
163
|
+
| `-o`, `--output` | Output file name to save results. | Yes | None |
|
|
164
|
+
|
|
165
|
+
### Example
|
|
166
|
+
|
|
167
|
+
To search for the top reported functions for taxon `g__Bacillus` in all `_Final_Contig.tsv` files within the `test_data/` directory:
|
|
168
|
+
|
|
169
|
+
```bash
|
|
170
|
+
Extract-By-Taxa -d .../test_data/Final_Contig -t g__Bacillus -o .../test_data/Final_Contig/Extract_By_Taxa/results.tsv -func GO
|
|
171
|
+
```
|
|
172
|
+
|
|
173
|
+
---
|
|
174
|
+
## Output
|
|
175
|
+
|
|
176
|
+
The tool generates a tab-delimited output file with the following columns:
|
|
177
|
+
|
|
178
|
+
1. **Sample:** Name of the processed Sample.
|
|
179
|
+
2. **Function:** Reported 'top' function.
|
|
180
|
+
3. **Num of Assignments (Functions):** Number of times the function has been assigned across all contigs reported as chosen Taxon.
|
|
181
|
+
|
|
182
|
+
Example output:
|
|
183
|
+
|
|
184
|
+
```
|
|
185
|
+
Selected Taxon: g__Bacillus
|
|
186
|
+
Sample Function Num of Assignments
|
|
187
|
+
PN0536_0001_S1 GO:0008150 296
|
|
188
|
+
PN0536_0001_S1 GO:0003674 285
|
|
189
|
+
PN0536_0001_S1 GO:0005575 254
|
|
190
|
+
PN0536_0003_S83 GO:0005575 45
|
|
191
|
+
PN0536_0003_S83 GO:0008150 44
|
|
192
|
+
PN0536_0003_S83 GO:0003674 43
|
|
193
|
+
PN0536_0002_S2 GO:0005575 5
|
|
194
|
+
PN0536_0002_S2 GO:0008150 5
|
|
195
|
+
PN0536_0002_S2 GO:0005623 4
|
|
196
|
+
PN0536_0004_S3 GO:0008150 4
|
|
197
|
+
PN0536_0004_S3 GO:0003674 3
|
|
198
|
+
PN0536_0004_S3 GO:0005488 3
|
|
199
|
+
|
|
200
|
+
```
|
|
201
|
+
|
|
202
|
+
---
|
|
203
|
+
|
|
204
|
+
### Large File Handling (Might be a failure point)
|
|
205
|
+
|
|
206
|
+
The script uses `csv.field_size_limit` to handle exceptionally large `.tsv` files.
|
|
207
|
+
|
|
208
|
+
---
|
|
209
|
+
|
|
210
|
+
|
metapont-0.0.3/README.md
ADDED
|
@@ -0,0 +1,195 @@
|
|
|
1
|
+
# MetaPont
|
|
2
|
+
**MetaPont** - A tool to bridge the gap between the output of metagenomic tools and the analysis of the data
|
|
3
|
+
|
|
4
|
+
## Features - These are the current aims of this project - Still under development
|
|
5
|
+
|
|
6
|
+
- **Targeted Functional Analysis:** Search for specific functional IDs (e.g., GO terms) within the `_Final_Contig.tsv` files provided by the HuwsLab Metagenome Workflow (https://github.com/TheHuwsLab/Metagenome_Workflow) .
|
|
7
|
+
- **Taxonomic Breakdown:** Extract genus-level taxonomy information and calculate their proportions in the dataset.
|
|
8
|
+
- **Batch Processing:** Analyse all `_Final_Contig.tsv` files in a specified directory.
|
|
9
|
+
- **Customisable Output:** Save results in a format suitable for downstream analysis.
|
|
10
|
+
|
|
11
|
+
---
|
|
12
|
+
|
|
13
|
+
## Installation
|
|
14
|
+
|
|
15
|
+
### Prerequisites
|
|
16
|
+
|
|
17
|
+
Ensure you have the following installed:
|
|
18
|
+
|
|
19
|
+
- Python ~3.6 or later
|
|
20
|
+
|
|
21
|
+
### Installation via pip
|
|
22
|
+
|
|
23
|
+
MetaPont is provided as a pip distribution.
|
|
24
|
+
|
|
25
|
+
```bash
|
|
26
|
+
pip install MetaPont
|
|
27
|
+
```
|
|
28
|
+
|
|
29
|
+
---
|
|
30
|
+
|
|
31
|
+
## Usage
|
|
32
|
+
|
|
33
|
+
---
|
|
34
|
+
### Extract-By-Function Command-line Arguments
|
|
35
|
+
```Extract-By-Function -h ```
|
|
36
|
+
```bash
|
|
37
|
+
usage: Extract_By_Function.py [-h] -d DIRECTORY -f FUNCTION_ID -o OUTPUT
|
|
38
|
+
[-m MIN_PROPORTION] [-top TOP_TAXA]
|
|
39
|
+
|
|
40
|
+
MetaPont v0.0.3: Extract-By-Function - Identify taxa contributing to a
|
|
41
|
+
specific function.
|
|
42
|
+
|
|
43
|
+
options:
|
|
44
|
+
-h, --help show this help message and exit
|
|
45
|
+
-d DIRECTORY, --directory DIRECTORY
|
|
46
|
+
Directory containing TSV files to analyse.
|
|
47
|
+
-f FUNCTION_ID, --function_id FUNCTION_ID
|
|
48
|
+
Specific function ID to search for (e.g.,
|
|
49
|
+
'GO:0016597').
|
|
50
|
+
-o OUTPUT, --output OUTPUT
|
|
51
|
+
Output file to save results.
|
|
52
|
+
-m MIN_PROPORTION, --min_proportion MIN_PROPORTION
|
|
53
|
+
Minimum proportion threshold for taxa to be included
|
|
54
|
+
in the output.
|
|
55
|
+
-top TOP_TAXA, --top_taxa TOP_TAXA
|
|
56
|
+
Top n taxa to be included in the output.
|
|
57
|
+
|
|
58
|
+
```
|
|
59
|
+
|
|
60
|
+
The `Extract-By-Function` tool provides several command-line options: \
|
|
61
|
+
Note: Either -m or -top is required.
|
|
62
|
+
|
|
63
|
+
| Option | Description | Required | Default |
|
|
64
|
+
|--------------------------|------------------------------------------------------------|----------|---------|
|
|
65
|
+
| `-d`, `--directory` | Directory containing `_Final_Contig.tsv` files to analyse. | Yes | None |
|
|
66
|
+
| `-f`, `--function_id` | Functional ID to search for (e.g., `GO:0016597`). | Yes | None |
|
|
67
|
+
| `-m`, `--min_proportion` | Minimum proportion needed for reporting. | Yes/No | None |
|
|
68
|
+
| `-top`, `--top_taxa` | Number of taxa to report. | Yes/No | None |
|
|
69
|
+
| `-o`, `--output` | Output file name to save results. | Yes | None |
|
|
70
|
+
|
|
71
|
+
### Example
|
|
72
|
+
|
|
73
|
+
To search for the functional ID `GO:0016597` in all `_Final_Contig.tsv` files within the `test_data/` directory:
|
|
74
|
+
|
|
75
|
+
```bash
|
|
76
|
+
Extract=By-Function -d .../test_data/Final_contig/ -f GO:0016597 -top 3 -o .../test_data/Final_Contig/Extract_By_Function_Out/results.tsv
|
|
77
|
+
```
|
|
78
|
+
|
|
79
|
+
---
|
|
80
|
+
|
|
81
|
+
## Output
|
|
82
|
+
|
|
83
|
+
The tool generates a tab-delimited output file with the following columns:
|
|
84
|
+
|
|
85
|
+
1. **Sample:** Name of the processed Sample.
|
|
86
|
+
2. **Taxa:** Genus-level taxonomic assignment extracted from the `Lineage` column.
|
|
87
|
+
3. **Reads Assigned (Function):** Number of reads assigned to contigs with the given functional ID.
|
|
88
|
+
3. **Proportion:** Proportion of reads assigned to contigs of stated Taxa with the given functional ID within the sample.
|
|
89
|
+
4. **Proportion (Total Reads):** Proportion of reads assigned to contigs of stated Taxa with the given functional ID within the total reads of the sample.
|
|
90
|
+
|
|
91
|
+
Example output:
|
|
92
|
+
|
|
93
|
+
```
|
|
94
|
+
Function ID: GO:0016597
|
|
95
|
+
Sample Taxa Reads Assigned (Function) Proportion (Function) Proportion (Total Reads)
|
|
96
|
+
PN0536_0001_S1_Final_Contig.tsv Lactobacillus 111963 0.602 0.004
|
|
97
|
+
PN0536_0003_S83_Final_Contig.tsv Lactobacillus 20072 0.457 0.001
|
|
98
|
+
PN0536_0002_S2_Final_Contig.tsv Acutalibacter 145222 0.795 0.005
|
|
99
|
+
PN0536_0004_S3_Final_Contig.tsv Lactobacillus 40076 0.404 0.002
|
|
100
|
+
```
|
|
101
|
+
|
|
102
|
+
---
|
|
103
|
+
|
|
104
|
+
|
|
105
|
+
|
|
106
|
+
### Workflow - unfinished
|
|
107
|
+
|
|
108
|
+
1. The script reads `_Final_Contig.tsv` files from the specified directory.
|
|
109
|
+
2. For each file, it searches for occurrences of the given functional ID within specific columns.
|
|
110
|
+
3. Matches are associated with genus-level taxonomic information extracted from the `Lineage` column.
|
|
111
|
+
4. Taxa proportions are calculated and saved to the output file.
|
|
112
|
+
|
|
113
|
+
---
|
|
114
|
+
## Extract-By-Taxa Command-line Arguments
|
|
115
|
+
```Extract-By-Taxa -h ```
|
|
116
|
+
```bash
|
|
117
|
+
usage: Extract_By_Taxa.py [-h] -d DIRECTORY -t TAXON -o OUTPUT -func
|
|
118
|
+
FUNCTIONAL_CLASSES [-top TOP_FUNCTIONS]
|
|
119
|
+
|
|
120
|
+
MetaPont: Extract Top Functions by Taxon
|
|
121
|
+
|
|
122
|
+
options:
|
|
123
|
+
-h, --help show this help message and exit
|
|
124
|
+
-d DIRECTORY, --directory DIRECTORY
|
|
125
|
+
Directory containing TSV files to analyse.
|
|
126
|
+
-t TAXON, --taxon TAXON
|
|
127
|
+
Target taxon to search for (e.g., 'g__Bacillus').
|
|
128
|
+
-o OUTPUT, --output OUTPUT
|
|
129
|
+
Output file to save results.
|
|
130
|
+
-func FUNCTIONAL_CLASSES, --functional_classes FUNCTIONAL_CLASSES
|
|
131
|
+
Which functional classes to report (e.g. GO,EC,KEGG
|
|
132
|
+
etc).
|
|
133
|
+
-top TOP_FUNCTIONS, --top_functions TOP_FUNCTIONS
|
|
134
|
+
Top n functions to include in the output for each
|
|
135
|
+
sample (default: 3).
|
|
136
|
+
|
|
137
|
+
```
|
|
138
|
+
|
|
139
|
+
The `Extract-By-Taxa` tool provides several command-line options:
|
|
140
|
+
|
|
141
|
+
|
|
142
|
+
| Option | Description | Required | Default |
|
|
143
|
+
|-----------------------------------|-------------------------------------------------------------|----------|---------|
|
|
144
|
+
| `-d`, `--directory` | Directory containing `_Fincal_Contig.tsv` files to analyse. | Yes | None |
|
|
145
|
+
| `-t`, `--taxon` | Taxa to search for (e.g., `g__Bacillus`). | Yes | None |
|
|
146
|
+
| `-func`, `--functional_classes` | Functional classes to report (e.g. GO,EC,KEGG etc). | Yes | None |
|
|
147
|
+
| `-top`, `--top_taxa` | Number of functions to report (default 3). | No | None |
|
|
148
|
+
| `-o`, `--output` | Output file name to save results. | Yes | None |
|
|
149
|
+
|
|
150
|
+
### Example
|
|
151
|
+
|
|
152
|
+
To search for the top reported functions for taxon `g__Bacillus` in all `_Final_Contig.tsv` files within the `test_data/` directory:
|
|
153
|
+
|
|
154
|
+
```bash
|
|
155
|
+
Extract-By-Taxa -d .../test_data/Final_Contig -t g__Bacillus -o .../test_data/Final_Contig/Extract_By_Taxa/results.tsv -func GO
|
|
156
|
+
```
|
|
157
|
+
|
|
158
|
+
---
|
|
159
|
+
## Output
|
|
160
|
+
|
|
161
|
+
The tool generates a tab-delimited output file with the following columns:
|
|
162
|
+
|
|
163
|
+
1. **Sample:** Name of the processed Sample.
|
|
164
|
+
2. **Function:** Reported 'top' function.
|
|
165
|
+
3. **Num of Assignments (Functions):** Number of times the function has been assigned across all contigs reported as chosen Taxon.
|
|
166
|
+
|
|
167
|
+
Example output:
|
|
168
|
+
|
|
169
|
+
```
|
|
170
|
+
Selected Taxon: g__Bacillus
|
|
171
|
+
Sample Function Num of Assignments
|
|
172
|
+
PN0536_0001_S1 GO:0008150 296
|
|
173
|
+
PN0536_0001_S1 GO:0003674 285
|
|
174
|
+
PN0536_0001_S1 GO:0005575 254
|
|
175
|
+
PN0536_0003_S83 GO:0005575 45
|
|
176
|
+
PN0536_0003_S83 GO:0008150 44
|
|
177
|
+
PN0536_0003_S83 GO:0003674 43
|
|
178
|
+
PN0536_0002_S2 GO:0005575 5
|
|
179
|
+
PN0536_0002_S2 GO:0008150 5
|
|
180
|
+
PN0536_0002_S2 GO:0005623 4
|
|
181
|
+
PN0536_0004_S3 GO:0008150 4
|
|
182
|
+
PN0536_0004_S3 GO:0003674 3
|
|
183
|
+
PN0536_0004_S3 GO:0005488 3
|
|
184
|
+
|
|
185
|
+
```
|
|
186
|
+
|
|
187
|
+
---
|
|
188
|
+
|
|
189
|
+
### Large File Handling (Might be a failure point)
|
|
190
|
+
|
|
191
|
+
The script uses `csv.field_size_limit` to handle exceptionally large `.tsv` files.
|
|
192
|
+
|
|
193
|
+
---
|
|
194
|
+
|
|
195
|
+
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
[metadata]
|
|
2
2
|
name = MetaPont
|
|
3
|
-
version = v0.0.
|
|
3
|
+
version = v0.0.3
|
|
4
4
|
author = Nicholas Dimonaco
|
|
5
5
|
author_email = nicholas@dimonaco.co.uk
|
|
6
6
|
description = MetaPont - A tool to bridge the gap between the output of metagenomic tools and the analysis of the data
|
|
@@ -28,6 +28,7 @@ include = *
|
|
|
28
28
|
[options.entry_points]
|
|
29
29
|
console_scripts =
|
|
30
30
|
Extract-By-Function = MetaPont.Extract_By_Function:main
|
|
31
|
+
Extract-By-Taxa = MetaPont.Extract_By_Taxa:main
|
|
31
32
|
|
|
32
33
|
[egg_info]
|
|
33
34
|
tag_build =
|
|
@@ -62,18 +62,26 @@ def main():
|
|
|
62
62
|
)
|
|
63
63
|
parser.add_argument(
|
|
64
64
|
"-f", "--function_id", required=True,
|
|
65
|
-
help="Specific function ID to search for (e.g., 'GO:
|
|
65
|
+
help="Specific function ID to search for (e.g., 'GO:0016597')."
|
|
66
66
|
)
|
|
67
67
|
parser.add_argument(
|
|
68
|
-
"-o", "--output",
|
|
69
|
-
help="Output file to save results
|
|
68
|
+
"-o", "--output", required=True,
|
|
69
|
+
help="Output file to save results."
|
|
70
70
|
)
|
|
71
71
|
parser.add_argument(
|
|
72
|
-
"-m", "--min_proportion", type=float,
|
|
73
|
-
help="Minimum proportion threshold for taxa to be included in the output
|
|
72
|
+
"-m", "--min_proportion", type=float,
|
|
73
|
+
help="Minimum proportion threshold for taxa to be included in the output."
|
|
74
|
+
)
|
|
75
|
+
parser.add_argument(
|
|
76
|
+
"-top", "--top_taxa", type=int,
|
|
77
|
+
help="Top n taxa to be included in the output."
|
|
74
78
|
)
|
|
75
79
|
|
|
76
80
|
options = parser.parse_args()
|
|
81
|
+
if not options.min_proportion and not options.top_taxa:
|
|
82
|
+
sys.exit("Error: Please specify either a minimum proportion or the number of top taxa to include in the output.")
|
|
83
|
+
elif options.min_proportion and options.top_taxa:
|
|
84
|
+
sys.exit("Error: Please specify either a minimum proportion or the number of top taxa to include in the output, not both.")
|
|
77
85
|
print("Running MetaPont: Extract-By-Function " + MetaPont_Version)
|
|
78
86
|
|
|
79
87
|
input_path = os.path.abspath(options.directory)
|
|
@@ -94,12 +102,21 @@ def main():
|
|
|
94
102
|
out.write("Function ID: " + options.function_id + "\n")
|
|
95
103
|
out.write("Sample\tTaxa\tReads Assigned (Function)\tProportion (Function)\tProportion (Total Reads)\n")
|
|
96
104
|
for sample, (taxa_reads_function, total_reads_function, total_reads_all) in all_results.items():
|
|
97
|
-
|
|
98
|
-
|
|
99
|
-
|
|
100
|
-
|
|
101
|
-
|
|
102
|
-
|
|
105
|
+
if options.min_proportion:
|
|
106
|
+
for taxa, reads_function in taxa_reads_function.items():
|
|
107
|
+
proportion_function = reads_function / total_reads_function if total_reads_function > 0 else 0
|
|
108
|
+
proportion_total = reads_function / total_reads_all if total_reads_all > 0 else 0
|
|
109
|
+
if proportion_function >= options.min_proportion: # Apply minimum proportion filter
|
|
110
|
+
out.write(f"{sample}\t{taxa}\t{reads_function}\t{proportion_function:.3f}\t{proportion_total:.3f}\n")
|
|
111
|
+
elif options.top_taxa:
|
|
112
|
+
sorted_taxa_reads = sorted(taxa_reads_function.items(), key=lambda x: x[1], reverse=True)
|
|
113
|
+
for i, (taxa, reads_function) in enumerate(sorted_taxa_reads):
|
|
114
|
+
if i < options.top_taxa:
|
|
115
|
+
proportion_function = reads_function / total_reads_function if total_reads_function > 0 else 0
|
|
116
|
+
proportion_total = reads_function / total_reads_all if total_reads_all > 0 else 0
|
|
117
|
+
out.write(f"{sample.replace('_Final_Contig.tsv','')}\t{taxa}\t{reads_function}\t{proportion_function:.3f}\t{proportion_total:.3f}\n")
|
|
118
|
+
else:
|
|
119
|
+
break
|
|
103
120
|
|
|
104
121
|
print(f"Results saved to {options.output}")
|
|
105
122
|
|
|
@@ -0,0 +1,130 @@
|
|
|
1
|
+
import argparse
|
|
2
|
+
import os
|
|
3
|
+
import csv
|
|
4
|
+
import sys
|
|
5
|
+
from collections import defaultdict
|
|
6
|
+
|
|
7
|
+
# Adjust constants import for module vs. standalone script usage
|
|
8
|
+
try:
|
|
9
|
+
from .constants import *
|
|
10
|
+
except (ModuleNotFoundError, ImportError, NameError, TypeError):
|
|
11
|
+
from constants import *
|
|
12
|
+
|
|
13
|
+
# Needed to account for the large CSV/TSV files
|
|
14
|
+
csv.field_size_limit(sys.maxsize)
|
|
15
|
+
|
|
16
|
+
|
|
17
|
+
def process_tsv_by_taxon(file_path, target_taxon):
|
|
18
|
+
"""
|
|
19
|
+
Processes a TSV file to calculate the top functions grouped by taxon.
|
|
20
|
+
|
|
21
|
+
Parameters:
|
|
22
|
+
file_path (str): Path to the TSV file.
|
|
23
|
+
target_taxon (str): Taxon to search for in the lineage column.
|
|
24
|
+
|
|
25
|
+
Returns:
|
|
26
|
+
taxon_function_reads (dict): A dictionary mapping functions to their total reads for the specified taxon.
|
|
27
|
+
total_reads_taxon (int): Total reads assigned to the specified taxon across all functions.
|
|
28
|
+
"""
|
|
29
|
+
taxon_function_presence = defaultdict(int) # Reads for each function for the specified taxon
|
|
30
|
+
total_reads_taxon = 0 # Total reads for the specified taxon
|
|
31
|
+
|
|
32
|
+
with open(file_path, "r") as tsv_file:
|
|
33
|
+
reader = csv.reader(tsv_file, delimiter="\t")
|
|
34
|
+
next(reader) # Skip the first row (sample name)
|
|
35
|
+
headers = next(reader) # Read headers from the second row
|
|
36
|
+
|
|
37
|
+
# Locate the necessary columns
|
|
38
|
+
taxa_idx = headers.index("Lineage")
|
|
39
|
+
reads_idx = headers.index("Mapped_Reads")
|
|
40
|
+
|
|
41
|
+
# Process each row
|
|
42
|
+
for idx, row in enumerate(reader):
|
|
43
|
+
if len(row) < len(headers):
|
|
44
|
+
continue # Skip malformed rows
|
|
45
|
+
|
|
46
|
+
lineage = row[taxa_idx]
|
|
47
|
+
reads = int(row[reads_idx]) # Get the number of reads for this row
|
|
48
|
+
|
|
49
|
+
# Check if this row matches the specified taxon
|
|
50
|
+
if target_taxon in lineage:
|
|
51
|
+
total_reads_taxon += reads # Increment total reads for the taxon
|
|
52
|
+
|
|
53
|
+
# Loop through the function columns (row[6:] onward)
|
|
54
|
+
for col_idx, cell in enumerate(row[6:], start=6):
|
|
55
|
+
if cell:
|
|
56
|
+
header = headers[col_idx]
|
|
57
|
+
if header not in taxon_function_presence:
|
|
58
|
+
taxon_function_presence[header] = defaultdict(int)
|
|
59
|
+
for function in cell.replace(',', '|').split('|'):
|
|
60
|
+
taxon_function_presence[header][function] += 1
|
|
61
|
+
|
|
62
|
+
return taxon_function_presence, total_reads_taxon
|
|
63
|
+
|
|
64
|
+
|
|
65
|
+
def main():
|
|
66
|
+
parser = argparse.ArgumentParser(description='MetaPont: Extract Top Functions by Taxon')
|
|
67
|
+
parser.add_argument(
|
|
68
|
+
"-d", "--directory", required=True,
|
|
69
|
+
help="Directory containing TSV files to analyse."
|
|
70
|
+
)
|
|
71
|
+
parser.add_argument(
|
|
72
|
+
"-t", "--taxon", required=True,
|
|
73
|
+
help="Target taxon to search for (e.g., 'g__Escherichia')."
|
|
74
|
+
)
|
|
75
|
+
parser.add_argument(
|
|
76
|
+
"-o", "--output", required=True,
|
|
77
|
+
help="Output file to save results."
|
|
78
|
+
)
|
|
79
|
+
parser.add_argument(
|
|
80
|
+
"-func", "--functional_classes", required=True,
|
|
81
|
+
help="Which functional classes to report (e.g. GO,EC,KEGG etc)."
|
|
82
|
+
)
|
|
83
|
+
parser.add_argument(
|
|
84
|
+
"-top", "--top_functions", type=int, default=3,
|
|
85
|
+
help="Top n functions to include in the output for each sample (default: 3)."
|
|
86
|
+
)
|
|
87
|
+
|
|
88
|
+
options = parser.parse_args()
|
|
89
|
+
print("Running MetaPont: Extract Top Functions by Taxon")
|
|
90
|
+
|
|
91
|
+
input_path = os.path.abspath(options.directory)
|
|
92
|
+
output_path = os.path.abspath(options.output)
|
|
93
|
+
|
|
94
|
+
all_results = {}
|
|
95
|
+
|
|
96
|
+
# Process each TSV file in the directory
|
|
97
|
+
for file_name in os.listdir(input_path):
|
|
98
|
+
if file_name.endswith("_Final_Contig.tsv"):
|
|
99
|
+
file_path = os.path.join(options.directory, file_name)
|
|
100
|
+
print(f"Processing file: {file_name}")
|
|
101
|
+
taxon_function_presence, total_reads_taxon = process_tsv_by_taxon(file_path, options.taxon)
|
|
102
|
+
all_results[file_name] = (taxon_function_presence, total_reads_taxon)
|
|
103
|
+
|
|
104
|
+
# Write results to output
|
|
105
|
+
with open(output_path, "w") as out:
|
|
106
|
+
out.write("Selected Taxon: " + options.taxon + "\n")
|
|
107
|
+
out.write("Sample\tFunction\tNum of Assignments\n")
|
|
108
|
+
for sample, (taxon_function_presence, total_assignments_taxon) in all_results.items():
|
|
109
|
+
for function, assignments_dict in taxon_function_presence.items():
|
|
110
|
+
if any(func_class in function for func_class in options.functional_classes.split(',')):
|
|
111
|
+
sorted_functions = sorted(assignments_dict.items(), key=lambda x: x[1], reverse=True)
|
|
112
|
+
for i, (sub_function, assignments) in enumerate(sorted_functions):
|
|
113
|
+
if i < options.top_functions:
|
|
114
|
+
#proportion_taxon = assignments / total_reads_taxon if total_reads_taxon > 0 else 0
|
|
115
|
+
out.write(f"{sample.replace('_Final_Contig.tsv', '')}\t{sub_function}\t{assignments}\n")
|
|
116
|
+
else:
|
|
117
|
+
break
|
|
118
|
+
# sorted_functions = sorted(taxon_function_presence.items(), key=lambda x: x[1], reverse=True)
|
|
119
|
+
# for i, (function, reads) in enumerate(sorted_functions):
|
|
120
|
+
# if i < options.top_functions:
|
|
121
|
+
# proportion_taxon = reads / total_reads_taxon if total_reads_taxon > 0 else 0
|
|
122
|
+
# out.write(f"{sample.replace('_Final_Contig.tsv', '')}\t{function}\t{reads}\t{proportion_taxon:.3f}\n")
|
|
123
|
+
# else:
|
|
124
|
+
# break
|
|
125
|
+
|
|
126
|
+
print(f"Results saved to {options.output}")
|
|
127
|
+
|
|
128
|
+
|
|
129
|
+
if __name__ == "__main__":
|
|
130
|
+
main()
|
|
@@ -0,0 +1,210 @@
|
|
|
1
|
+
Metadata-Version: 2.1
|
|
2
|
+
Name: MetaPont
|
|
3
|
+
Version: 0.0.3
|
|
4
|
+
Summary: MetaPont - A tool to bridge the gap between the output of metagenomic tools and the analysis of the data
|
|
5
|
+
Home-page: https://github.com/TheHuwsLab/MetaPont
|
|
6
|
+
Author: Nicholas Dimonaco
|
|
7
|
+
Author-email: nicholas@dimonaco.co.uk
|
|
8
|
+
Project-URL: Bug Tracker, https://github.com/TheHuwsLab/MetaPont/issues
|
|
9
|
+
Classifier: Programming Language :: Python :: 3
|
|
10
|
+
Classifier: License :: OSI Approved :: GNU General Public License v3 (GPLv3)
|
|
11
|
+
Classifier: Operating System :: OS Independent
|
|
12
|
+
Requires-Python: >=3.6
|
|
13
|
+
Description-Content-Type: text/markdown
|
|
14
|
+
License-File: LICENSE
|
|
15
|
+
|
|
16
|
+
# MetaPont
|
|
17
|
+
**MetaPont** - A tool to bridge the gap between the output of metagenomic tools and the analysis of the data
|
|
18
|
+
|
|
19
|
+
## Features - These are the current aims of this project - Still under development
|
|
20
|
+
|
|
21
|
+
- **Targeted Functional Analysis:** Search for specific functional IDs (e.g., GO terms) within the `_Final_Contig.tsv` files provided by the HuwsLab Metagenome Workflow (https://github.com/TheHuwsLab/Metagenome_Workflow) .
|
|
22
|
+
- **Taxonomic Breakdown:** Extract genus-level taxonomy information and calculate their proportions in the dataset.
|
|
23
|
+
- **Batch Processing:** Analyse all `_Final_Contig.tsv` files in a specified directory.
|
|
24
|
+
- **Customisable Output:** Save results in a format suitable for downstream analysis.
|
|
25
|
+
|
|
26
|
+
---
|
|
27
|
+
|
|
28
|
+
## Installation
|
|
29
|
+
|
|
30
|
+
### Prerequisites
|
|
31
|
+
|
|
32
|
+
Ensure you have the following installed:
|
|
33
|
+
|
|
34
|
+
- Python ~3.6 or later
|
|
35
|
+
|
|
36
|
+
### Installation via pip
|
|
37
|
+
|
|
38
|
+
MetaPont is provided as a pip distribution.
|
|
39
|
+
|
|
40
|
+
```bash
|
|
41
|
+
pip install MetaPont
|
|
42
|
+
```
|
|
43
|
+
|
|
44
|
+
---
|
|
45
|
+
|
|
46
|
+
## Usage
|
|
47
|
+
|
|
48
|
+
---
|
|
49
|
+
### Extract-By-Function Command-line Arguments
|
|
50
|
+
```Extract-By-Function -h ```
|
|
51
|
+
```bash
|
|
52
|
+
usage: Extract_By_Function.py [-h] -d DIRECTORY -f FUNCTION_ID -o OUTPUT
|
|
53
|
+
[-m MIN_PROPORTION] [-top TOP_TAXA]
|
|
54
|
+
|
|
55
|
+
MetaPont v0.0.3: Extract-By-Function - Identify taxa contributing to a
|
|
56
|
+
specific function.
|
|
57
|
+
|
|
58
|
+
options:
|
|
59
|
+
-h, --help show this help message and exit
|
|
60
|
+
-d DIRECTORY, --directory DIRECTORY
|
|
61
|
+
Directory containing TSV files to analyse.
|
|
62
|
+
-f FUNCTION_ID, --function_id FUNCTION_ID
|
|
63
|
+
Specific function ID to search for (e.g.,
|
|
64
|
+
'GO:0016597').
|
|
65
|
+
-o OUTPUT, --output OUTPUT
|
|
66
|
+
Output file to save results.
|
|
67
|
+
-m MIN_PROPORTION, --min_proportion MIN_PROPORTION
|
|
68
|
+
Minimum proportion threshold for taxa to be included
|
|
69
|
+
in the output.
|
|
70
|
+
-top TOP_TAXA, --top_taxa TOP_TAXA
|
|
71
|
+
Top n taxa to be included in the output.
|
|
72
|
+
|
|
73
|
+
```
|
|
74
|
+
|
|
75
|
+
The `Extract-By-Function` tool provides several command-line options: \
|
|
76
|
+
Note: Either -m or -top is required.
|
|
77
|
+
|
|
78
|
+
| Option | Description | Required | Default |
|
|
79
|
+
|--------------------------|------------------------------------------------------------|----------|---------|
|
|
80
|
+
| `-d`, `--directory` | Directory containing `_Final_Contig.tsv` files to analyse. | Yes | None |
|
|
81
|
+
| `-f`, `--function_id` | Functional ID to search for (e.g., `GO:0016597`). | Yes | None |
|
|
82
|
+
| `-m`, `--min_proportion` | Minimum proportion needed for reporting. | Yes/No | None |
|
|
83
|
+
| `-top`, `--top_taxa` | Number of taxa to report. | Yes/No | None |
|
|
84
|
+
| `-o`, `--output` | Output file name to save results. | Yes | None |
|
|
85
|
+
|
|
86
|
+
### Example
|
|
87
|
+
|
|
88
|
+
To search for the functional ID `GO:0016597` in all `_Final_Contig.tsv` files within the `test_data/` directory:
|
|
89
|
+
|
|
90
|
+
```bash
|
|
91
|
+
Extract=By-Function -d .../test_data/Final_contig/ -f GO:0016597 -top 3 -o .../test_data/Final_Contig/Extract_By_Function_Out/results.tsv
|
|
92
|
+
```
|
|
93
|
+
|
|
94
|
+
---
|
|
95
|
+
|
|
96
|
+
## Output
|
|
97
|
+
|
|
98
|
+
The tool generates a tab-delimited output file with the following columns:
|
|
99
|
+
|
|
100
|
+
1. **Sample:** Name of the processed Sample.
|
|
101
|
+
2. **Taxa:** Genus-level taxonomic assignment extracted from the `Lineage` column.
|
|
102
|
+
3. **Reads Assigned (Function):** Number of reads assigned to contigs with the given functional ID.
|
|
103
|
+
3. **Proportion:** Proportion of reads assigned to contigs of stated Taxa with the given functional ID within the sample.
|
|
104
|
+
4. **Proportion (Total Reads):** Proportion of reads assigned to contigs of stated Taxa with the given functional ID within the total reads of the sample.
|
|
105
|
+
|
|
106
|
+
Example output:
|
|
107
|
+
|
|
108
|
+
```
|
|
109
|
+
Function ID: GO:0016597
|
|
110
|
+
Sample Taxa Reads Assigned (Function) Proportion (Function) Proportion (Total Reads)
|
|
111
|
+
PN0536_0001_S1_Final_Contig.tsv Lactobacillus 111963 0.602 0.004
|
|
112
|
+
PN0536_0003_S83_Final_Contig.tsv Lactobacillus 20072 0.457 0.001
|
|
113
|
+
PN0536_0002_S2_Final_Contig.tsv Acutalibacter 145222 0.795 0.005
|
|
114
|
+
PN0536_0004_S3_Final_Contig.tsv Lactobacillus 40076 0.404 0.002
|
|
115
|
+
```
|
|
116
|
+
|
|
117
|
+
---
|
|
118
|
+
|
|
119
|
+
|
|
120
|
+
|
|
121
|
+
### Workflow - unfinished
|
|
122
|
+
|
|
123
|
+
1. The script reads `_Final_Contig.tsv` files from the specified directory.
|
|
124
|
+
2. For each file, it searches for occurrences of the given functional ID within specific columns.
|
|
125
|
+
3. Matches are associated with genus-level taxonomic information extracted from the `Lineage` column.
|
|
126
|
+
4. Taxa proportions are calculated and saved to the output file.
|
|
127
|
+
|
|
128
|
+
---
|
|
129
|
+
## Extract-By-Taxa Command-line Arguments
|
|
130
|
+
```Extract-By-Taxa -h ```
|
|
131
|
+
```bash
|
|
132
|
+
usage: Extract_By_Taxa.py [-h] -d DIRECTORY -t TAXON -o OUTPUT -func
|
|
133
|
+
FUNCTIONAL_CLASSES [-top TOP_FUNCTIONS]
|
|
134
|
+
|
|
135
|
+
MetaPont: Extract Top Functions by Taxon
|
|
136
|
+
|
|
137
|
+
options:
|
|
138
|
+
-h, --help show this help message and exit
|
|
139
|
+
-d DIRECTORY, --directory DIRECTORY
|
|
140
|
+
Directory containing TSV files to analyse.
|
|
141
|
+
-t TAXON, --taxon TAXON
|
|
142
|
+
Target taxon to search for (e.g., 'g__Bacillus').
|
|
143
|
+
-o OUTPUT, --output OUTPUT
|
|
144
|
+
Output file to save results.
|
|
145
|
+
-func FUNCTIONAL_CLASSES, --functional_classes FUNCTIONAL_CLASSES
|
|
146
|
+
Which functional classes to report (e.g. GO,EC,KEGG
|
|
147
|
+
etc).
|
|
148
|
+
-top TOP_FUNCTIONS, --top_functions TOP_FUNCTIONS
|
|
149
|
+
Top n functions to include in the output for each
|
|
150
|
+
sample (default: 3).
|
|
151
|
+
|
|
152
|
+
```
|
|
153
|
+
|
|
154
|
+
The `Extract-By-Taxa` tool provides several command-line options:
|
|
155
|
+
|
|
156
|
+
|
|
157
|
+
| Option | Description | Required | Default |
|
|
158
|
+
|-----------------------------------|-------------------------------------------------------------|----------|---------|
|
|
159
|
+
| `-d`, `--directory` | Directory containing `_Fincal_Contig.tsv` files to analyse. | Yes | None |
|
|
160
|
+
| `-t`, `--taxon` | Taxa to search for (e.g., `g__Bacillus`). | Yes | None |
|
|
161
|
+
| `-func`, `--functional_classes` | Functional classes to report (e.g. GO,EC,KEGG etc). | Yes | None |
|
|
162
|
+
| `-top`, `--top_taxa` | Number of functions to report (default 3). | No | None |
|
|
163
|
+
| `-o`, `--output` | Output file name to save results. | Yes | None |
|
|
164
|
+
|
|
165
|
+
### Example
|
|
166
|
+
|
|
167
|
+
To search for the top reported functions for taxon `g__Bacillus` in all `_Final_Contig.tsv` files within the `test_data/` directory:
|
|
168
|
+
|
|
169
|
+
```bash
|
|
170
|
+
Extract-By-Taxa -d .../test_data/Final_Contig -t g__Bacillus -o .../test_data/Final_Contig/Extract_By_Taxa/results.tsv -func GO
|
|
171
|
+
```
|
|
172
|
+
|
|
173
|
+
---
|
|
174
|
+
## Output
|
|
175
|
+
|
|
176
|
+
The tool generates a tab-delimited output file with the following columns:
|
|
177
|
+
|
|
178
|
+
1. **Sample:** Name of the processed Sample.
|
|
179
|
+
2. **Function:** Reported 'top' function.
|
|
180
|
+
3. **Num of Assignments (Functions):** Number of times the function has been assigned across all contigs reported as chosen Taxon.
|
|
181
|
+
|
|
182
|
+
Example output:
|
|
183
|
+
|
|
184
|
+
```
|
|
185
|
+
Selected Taxon: g__Bacillus
|
|
186
|
+
Sample Function Num of Assignments
|
|
187
|
+
PN0536_0001_S1 GO:0008150 296
|
|
188
|
+
PN0536_0001_S1 GO:0003674 285
|
|
189
|
+
PN0536_0001_S1 GO:0005575 254
|
|
190
|
+
PN0536_0003_S83 GO:0005575 45
|
|
191
|
+
PN0536_0003_S83 GO:0008150 44
|
|
192
|
+
PN0536_0003_S83 GO:0003674 43
|
|
193
|
+
PN0536_0002_S2 GO:0005575 5
|
|
194
|
+
PN0536_0002_S2 GO:0008150 5
|
|
195
|
+
PN0536_0002_S2 GO:0005623 4
|
|
196
|
+
PN0536_0004_S3 GO:0008150 4
|
|
197
|
+
PN0536_0004_S3 GO:0003674 3
|
|
198
|
+
PN0536_0004_S3 GO:0005488 3
|
|
199
|
+
|
|
200
|
+
```
|
|
201
|
+
|
|
202
|
+
---
|
|
203
|
+
|
|
204
|
+
### Large File Handling (Might be a failure point)
|
|
205
|
+
|
|
206
|
+
The script uses `csv.field_size_limit` to handle exceptionally large `.tsv` files.
|
|
207
|
+
|
|
208
|
+
---
|
|
209
|
+
|
|
210
|
+
|
metapont-0.0.2/PKG-INFO
DELETED
|
@@ -1,130 +0,0 @@
|
|
|
1
|
-
Metadata-Version: 2.1
|
|
2
|
-
Name: MetaPont
|
|
3
|
-
Version: 0.0.2
|
|
4
|
-
Summary: MetaPont - A tool to bridge the gap between the output of metagenomic tools and the analysis of the data
|
|
5
|
-
Home-page: https://github.com/TheHuwsLab/MetaPont
|
|
6
|
-
Author: Nicholas Dimonaco
|
|
7
|
-
Author-email: nicholas@dimonaco.co.uk
|
|
8
|
-
Project-URL: Bug Tracker, https://github.com/TheHuwsLab/MetaPont/issues
|
|
9
|
-
Classifier: Programming Language :: Python :: 3
|
|
10
|
-
Classifier: License :: OSI Approved :: GNU General Public License v3 (GPLv3)
|
|
11
|
-
Classifier: Operating System :: OS Independent
|
|
12
|
-
Requires-Python: >=3.6
|
|
13
|
-
Description-Content-Type: text/markdown
|
|
14
|
-
License-File: LICENSE
|
|
15
|
-
|
|
16
|
-
# MetaPont
|
|
17
|
-
**MetaPont** - A tool to bridge the gap between the output of metagenomic tools and the analysis of the data
|
|
18
|
-
|
|
19
|
-
## Features - These are the current aims of this project - Still under development
|
|
20
|
-
|
|
21
|
-
- **Targeted Functional Analysis:** Search for specific functional IDs (e.g., GO terms) within the `.tsv` files provided by the HuwsLab Metagenome Workflow (https://github.com/TheHuwsLab/Metagenome_Workflow) .
|
|
22
|
-
- **Taxonomic Breakdown:** Extract genus-level taxonomy information and calculate their proportions in the dataset.
|
|
23
|
-
- **Batch Processing:** Analyse all `.tsv` files in a specified directory.
|
|
24
|
-
- **Customisable Output:** Save results in a format suitable for downstream analysis.
|
|
25
|
-
|
|
26
|
-
---
|
|
27
|
-
|
|
28
|
-
## Installation
|
|
29
|
-
|
|
30
|
-
### Prerequisites
|
|
31
|
-
|
|
32
|
-
Ensure you have the following installed:
|
|
33
|
-
|
|
34
|
-
- Python ~3.6 or later
|
|
35
|
-
- Required Python libraries: `argparse`, `csv`, and `collections` (standard libs).
|
|
36
|
-
|
|
37
|
-
### Installation via pip
|
|
38
|
-
|
|
39
|
-
MetaPont is provided as a pip distribution.
|
|
40
|
-
|
|
41
|
-
```bash
|
|
42
|
-
pip install MetaPont
|
|
43
|
-
```
|
|
44
|
-
|
|
45
|
-
---
|
|
46
|
-
|
|
47
|
-
## Usage
|
|
48
|
-
|
|
49
|
-
### Command-line Arguments
|
|
50
|
-
```Extract-By-Function -h ```
|
|
51
|
-
```bash
|
|
52
|
-
usage: Extract-By-Function [-h] -d DIRECTORY -f FUNCTION_ID [-o OUTPUT] [-m MIN_PROPORTION]
|
|
53
|
-
|
|
54
|
-
MetaPont v0.0.2: Extract-By-Function - Identify taxa contributing to a specific function.
|
|
55
|
-
|
|
56
|
-
options:
|
|
57
|
-
-h, --help show this help message and exit
|
|
58
|
-
-d DIRECTORY, --directory DIRECTORY
|
|
59
|
-
Directory containing TSV files to analyse.
|
|
60
|
-
-f FUNCTION_ID, --function_id FUNCTION_ID
|
|
61
|
-
Specific function ID to search for (e.g., 'GO:0002').
|
|
62
|
-
-o OUTPUT, --output OUTPUT
|
|
63
|
-
Output file to save results (default: output_taxa_details.tsv).
|
|
64
|
-
-m MIN_PROPORTION, --min_proportion MIN_PROPORTION
|
|
65
|
-
Minimum proportion threshold for taxa to be included in the output (default: 0.05).
|
|
66
|
-
```
|
|
67
|
-
|
|
68
|
-
The `Extract-By-Function` tool provides several command-line options:
|
|
69
|
-
|
|
70
|
-
| Option | Description | Required | Default |
|
|
71
|
-
|--------------------------|-----------------------------------------------|----------|-------------------------------|
|
|
72
|
-
| `-d`, `--directory` | Directory containing `.tsv` files to analyse. | Yes | None |
|
|
73
|
-
| `-f`, `--function_id` | Functional ID to search for (e.g., `GO:0002`). | Yes | None |
|
|
74
|
-
| `-m`, `--min_proportion` | Minimum proportion needed for reporting. | Yes | 0.05 (5%) |
|
|
75
|
-
| `-o`, `--output` | Output file name to save results. | No | `output_taxa_proportions.tsv` |
|
|
76
|
-
|
|
77
|
-
### Example
|
|
78
|
-
|
|
79
|
-
To search for the functional ID `GO:0002` in all `.tsv` files within the `data/` directory:
|
|
80
|
-
|
|
81
|
-
```bash
|
|
82
|
-
ExtractByFunction -d .../test_data/Final_contig/ -f GO:0002 -m 0.10 -o .../test_data/Final_Contig/Extract_By_Function_Out/results.tsv
|
|
83
|
-
```
|
|
84
|
-
|
|
85
|
-
---
|
|
86
|
-
|
|
87
|
-
## Output
|
|
88
|
-
|
|
89
|
-
The tool generates a tab-delimited output file with the following columns:
|
|
90
|
-
|
|
91
|
-
1. **Sample:** Name of the processed `.tsv` file.
|
|
92
|
-
2. **Taxa:** Genus-level taxonomic assignment extracted from the `Lineage` column.
|
|
93
|
-
3. **Proportion:** Proportion of matches to the given functional ID within the sample.
|
|
94
|
-
|
|
95
|
-
Example output:
|
|
96
|
-
|
|
97
|
-
```
|
|
98
|
-
Function ID: GO:0002
|
|
99
|
-
Sample Taxa Reads Assigned (Function) Proportion (Function) Proportion (Total Reads)
|
|
100
|
-
PN0536_0003_S83_Final_Contig.tsv Gordonibacter 60788 0.075 0.002
|
|
101
|
-
PN0536_0003_S83_Final_Contig.tsv Streptomyces 115671 0.142 0.004
|
|
102
|
-
PN0536_0003_S83_Final_Contig.tsv unknown 80890 0.099 0.003
|
|
103
|
-
PN0536_0003_S83_Final_Contig.tsv Clostridium 51018 0.063 0.002
|
|
104
|
-
PN0536_0003_S83_Final_Contig.tsv Lactobacillus 149909 0.184 0.005
|
|
105
|
-
PN0536_0003_S83_Final_Contig.tsv Limosilactobacillus 79694 0.098 0.003
|
|
106
|
-
```
|
|
107
|
-
|
|
108
|
-
---
|
|
109
|
-
|
|
110
|
-
## Implementation Details
|
|
111
|
-
|
|
112
|
-
### Workflow
|
|
113
|
-
|
|
114
|
-
1. The script reads `.tsv` files from the specified directory.
|
|
115
|
-
2. For each file, it searches for occurrences of the given functional ID within specific columns.
|
|
116
|
-
3. Matches are associated with genus-level taxonomic information extracted from the `Lineage` column.
|
|
117
|
-
4. Taxa proportions are calculated and saved to the output file.
|
|
118
|
-
|
|
119
|
-
### Large File Handling (Might be a failure point)
|
|
120
|
-
|
|
121
|
-
The script uses `csv.field_size_limit` to handle exceptionally large `.tsv` files.
|
|
122
|
-
|
|
123
|
-
---
|
|
124
|
-
|
|
125
|
-
## Future Plans
|
|
126
|
-
|
|
127
|
-
- Add support for additional file formats (e.g., `.csv`, `.txt`).
|
|
128
|
-
- Expand functionality for more complex taxonomic and functional analyses.
|
|
129
|
-
---
|
|
130
|
-
|
metapont-0.0.2/README.md
DELETED
|
@@ -1,115 +0,0 @@
|
|
|
1
|
-
# MetaPont
|
|
2
|
-
**MetaPont** - A tool to bridge the gap between the output of metagenomic tools and the analysis of the data
|
|
3
|
-
|
|
4
|
-
## Features - These are the current aims of this project - Still under development
|
|
5
|
-
|
|
6
|
-
- **Targeted Functional Analysis:** Search for specific functional IDs (e.g., GO terms) within the `.tsv` files provided by the HuwsLab Metagenome Workflow (https://github.com/TheHuwsLab/Metagenome_Workflow) .
|
|
7
|
-
- **Taxonomic Breakdown:** Extract genus-level taxonomy information and calculate their proportions in the dataset.
|
|
8
|
-
- **Batch Processing:** Analyse all `.tsv` files in a specified directory.
|
|
9
|
-
- **Customisable Output:** Save results in a format suitable for downstream analysis.
|
|
10
|
-
|
|
11
|
-
---
|
|
12
|
-
|
|
13
|
-
## Installation
|
|
14
|
-
|
|
15
|
-
### Prerequisites
|
|
16
|
-
|
|
17
|
-
Ensure you have the following installed:
|
|
18
|
-
|
|
19
|
-
- Python ~3.6 or later
|
|
20
|
-
- Required Python libraries: `argparse`, `csv`, and `collections` (standard libs).
|
|
21
|
-
|
|
22
|
-
### Installation via pip
|
|
23
|
-
|
|
24
|
-
MetaPont is provided as a pip distribution.
|
|
25
|
-
|
|
26
|
-
```bash
|
|
27
|
-
pip install MetaPont
|
|
28
|
-
```
|
|
29
|
-
|
|
30
|
-
---
|
|
31
|
-
|
|
32
|
-
## Usage
|
|
33
|
-
|
|
34
|
-
### Command-line Arguments
|
|
35
|
-
```Extract-By-Function -h ```
|
|
36
|
-
```bash
|
|
37
|
-
usage: Extract-By-Function [-h] -d DIRECTORY -f FUNCTION_ID [-o OUTPUT] [-m MIN_PROPORTION]
|
|
38
|
-
|
|
39
|
-
MetaPont v0.0.2: Extract-By-Function - Identify taxa contributing to a specific function.
|
|
40
|
-
|
|
41
|
-
options:
|
|
42
|
-
-h, --help show this help message and exit
|
|
43
|
-
-d DIRECTORY, --directory DIRECTORY
|
|
44
|
-
Directory containing TSV files to analyse.
|
|
45
|
-
-f FUNCTION_ID, --function_id FUNCTION_ID
|
|
46
|
-
Specific function ID to search for (e.g., 'GO:0002').
|
|
47
|
-
-o OUTPUT, --output OUTPUT
|
|
48
|
-
Output file to save results (default: output_taxa_details.tsv).
|
|
49
|
-
-m MIN_PROPORTION, --min_proportion MIN_PROPORTION
|
|
50
|
-
Minimum proportion threshold for taxa to be included in the output (default: 0.05).
|
|
51
|
-
```
|
|
52
|
-
|
|
53
|
-
The `Extract-By-Function` tool provides several command-line options:
|
|
54
|
-
|
|
55
|
-
| Option | Description | Required | Default |
|
|
56
|
-
|--------------------------|-----------------------------------------------|----------|-------------------------------|
|
|
57
|
-
| `-d`, `--directory` | Directory containing `.tsv` files to analyse. | Yes | None |
|
|
58
|
-
| `-f`, `--function_id` | Functional ID to search for (e.g., `GO:0002`). | Yes | None |
|
|
59
|
-
| `-m`, `--min_proportion` | Minimum proportion needed for reporting. | Yes | 0.05 (5%) |
|
|
60
|
-
| `-o`, `--output` | Output file name to save results. | No | `output_taxa_proportions.tsv` |
|
|
61
|
-
|
|
62
|
-
### Example
|
|
63
|
-
|
|
64
|
-
To search for the functional ID `GO:0002` in all `.tsv` files within the `data/` directory:
|
|
65
|
-
|
|
66
|
-
```bash
|
|
67
|
-
ExtractByFunction -d .../test_data/Final_contig/ -f GO:0002 -m 0.10 -o .../test_data/Final_Contig/Extract_By_Function_Out/results.tsv
|
|
68
|
-
```
|
|
69
|
-
|
|
70
|
-
---
|
|
71
|
-
|
|
72
|
-
## Output
|
|
73
|
-
|
|
74
|
-
The tool generates a tab-delimited output file with the following columns:
|
|
75
|
-
|
|
76
|
-
1. **Sample:** Name of the processed `.tsv` file.
|
|
77
|
-
2. **Taxa:** Genus-level taxonomic assignment extracted from the `Lineage` column.
|
|
78
|
-
3. **Proportion:** Proportion of matches to the given functional ID within the sample.
|
|
79
|
-
|
|
80
|
-
Example output:
|
|
81
|
-
|
|
82
|
-
```
|
|
83
|
-
Function ID: GO:0002
|
|
84
|
-
Sample Taxa Reads Assigned (Function) Proportion (Function) Proportion (Total Reads)
|
|
85
|
-
PN0536_0003_S83_Final_Contig.tsv Gordonibacter 60788 0.075 0.002
|
|
86
|
-
PN0536_0003_S83_Final_Contig.tsv Streptomyces 115671 0.142 0.004
|
|
87
|
-
PN0536_0003_S83_Final_Contig.tsv unknown 80890 0.099 0.003
|
|
88
|
-
PN0536_0003_S83_Final_Contig.tsv Clostridium 51018 0.063 0.002
|
|
89
|
-
PN0536_0003_S83_Final_Contig.tsv Lactobacillus 149909 0.184 0.005
|
|
90
|
-
PN0536_0003_S83_Final_Contig.tsv Limosilactobacillus 79694 0.098 0.003
|
|
91
|
-
```
|
|
92
|
-
|
|
93
|
-
---
|
|
94
|
-
|
|
95
|
-
## Implementation Details
|
|
96
|
-
|
|
97
|
-
### Workflow
|
|
98
|
-
|
|
99
|
-
1. The script reads `.tsv` files from the specified directory.
|
|
100
|
-
2. For each file, it searches for occurrences of the given functional ID within specific columns.
|
|
101
|
-
3. Matches are associated with genus-level taxonomic information extracted from the `Lineage` column.
|
|
102
|
-
4. Taxa proportions are calculated and saved to the output file.
|
|
103
|
-
|
|
104
|
-
### Large File Handling (Might be a failure point)
|
|
105
|
-
|
|
106
|
-
The script uses `csv.field_size_limit` to handle exceptionally large `.tsv` files.
|
|
107
|
-
|
|
108
|
-
---
|
|
109
|
-
|
|
110
|
-
## Future Plans
|
|
111
|
-
|
|
112
|
-
- Add support for additional file formats (e.g., `.csv`, `.txt`).
|
|
113
|
-
- Expand functionality for more complex taxonomic and functional analyses.
|
|
114
|
-
---
|
|
115
|
-
|
|
@@ -1,130 +0,0 @@
|
|
|
1
|
-
Metadata-Version: 2.1
|
|
2
|
-
Name: MetaPont
|
|
3
|
-
Version: 0.0.2
|
|
4
|
-
Summary: MetaPont - A tool to bridge the gap between the output of metagenomic tools and the analysis of the data
|
|
5
|
-
Home-page: https://github.com/TheHuwsLab/MetaPont
|
|
6
|
-
Author: Nicholas Dimonaco
|
|
7
|
-
Author-email: nicholas@dimonaco.co.uk
|
|
8
|
-
Project-URL: Bug Tracker, https://github.com/TheHuwsLab/MetaPont/issues
|
|
9
|
-
Classifier: Programming Language :: Python :: 3
|
|
10
|
-
Classifier: License :: OSI Approved :: GNU General Public License v3 (GPLv3)
|
|
11
|
-
Classifier: Operating System :: OS Independent
|
|
12
|
-
Requires-Python: >=3.6
|
|
13
|
-
Description-Content-Type: text/markdown
|
|
14
|
-
License-File: LICENSE
|
|
15
|
-
|
|
16
|
-
# MetaPont
|
|
17
|
-
**MetaPont** - A tool to bridge the gap between the output of metagenomic tools and the analysis of the data
|
|
18
|
-
|
|
19
|
-
## Features - These are the current aims of this project - Still under development
|
|
20
|
-
|
|
21
|
-
- **Targeted Functional Analysis:** Search for specific functional IDs (e.g., GO terms) within the `.tsv` files provided by the HuwsLab Metagenome Workflow (https://github.com/TheHuwsLab/Metagenome_Workflow) .
|
|
22
|
-
- **Taxonomic Breakdown:** Extract genus-level taxonomy information and calculate their proportions in the dataset.
|
|
23
|
-
- **Batch Processing:** Analyse all `.tsv` files in a specified directory.
|
|
24
|
-
- **Customisable Output:** Save results in a format suitable for downstream analysis.
|
|
25
|
-
|
|
26
|
-
---
|
|
27
|
-
|
|
28
|
-
## Installation
|
|
29
|
-
|
|
30
|
-
### Prerequisites
|
|
31
|
-
|
|
32
|
-
Ensure you have the following installed:
|
|
33
|
-
|
|
34
|
-
- Python ~3.6 or later
|
|
35
|
-
- Required Python libraries: `argparse`, `csv`, and `collections` (standard libs).
|
|
36
|
-
|
|
37
|
-
### Installation via pip
|
|
38
|
-
|
|
39
|
-
MetaPont is provided as a pip distribution.
|
|
40
|
-
|
|
41
|
-
```bash
|
|
42
|
-
pip install MetaPont
|
|
43
|
-
```
|
|
44
|
-
|
|
45
|
-
---
|
|
46
|
-
|
|
47
|
-
## Usage
|
|
48
|
-
|
|
49
|
-
### Command-line Arguments
|
|
50
|
-
```Extract-By-Function -h ```
|
|
51
|
-
```bash
|
|
52
|
-
usage: Extract-By-Function [-h] -d DIRECTORY -f FUNCTION_ID [-o OUTPUT] [-m MIN_PROPORTION]
|
|
53
|
-
|
|
54
|
-
MetaPont v0.0.2: Extract-By-Function - Identify taxa contributing to a specific function.
|
|
55
|
-
|
|
56
|
-
options:
|
|
57
|
-
-h, --help show this help message and exit
|
|
58
|
-
-d DIRECTORY, --directory DIRECTORY
|
|
59
|
-
Directory containing TSV files to analyse.
|
|
60
|
-
-f FUNCTION_ID, --function_id FUNCTION_ID
|
|
61
|
-
Specific function ID to search for (e.g., 'GO:0002').
|
|
62
|
-
-o OUTPUT, --output OUTPUT
|
|
63
|
-
Output file to save results (default: output_taxa_details.tsv).
|
|
64
|
-
-m MIN_PROPORTION, --min_proportion MIN_PROPORTION
|
|
65
|
-
Minimum proportion threshold for taxa to be included in the output (default: 0.05).
|
|
66
|
-
```
|
|
67
|
-
|
|
68
|
-
The `Extract-By-Function` tool provides several command-line options:
|
|
69
|
-
|
|
70
|
-
| Option | Description | Required | Default |
|
|
71
|
-
|--------------------------|-----------------------------------------------|----------|-------------------------------|
|
|
72
|
-
| `-d`, `--directory` | Directory containing `.tsv` files to analyse. | Yes | None |
|
|
73
|
-
| `-f`, `--function_id` | Functional ID to search for (e.g., `GO:0002`). | Yes | None |
|
|
74
|
-
| `-m`, `--min_proportion` | Minimum proportion needed for reporting. | Yes | 0.05 (5%) |
|
|
75
|
-
| `-o`, `--output` | Output file name to save results. | No | `output_taxa_proportions.tsv` |
|
|
76
|
-
|
|
77
|
-
### Example
|
|
78
|
-
|
|
79
|
-
To search for the functional ID `GO:0002` in all `.tsv` files within the `data/` directory:
|
|
80
|
-
|
|
81
|
-
```bash
|
|
82
|
-
ExtractByFunction -d .../test_data/Final_contig/ -f GO:0002 -m 0.10 -o .../test_data/Final_Contig/Extract_By_Function_Out/results.tsv
|
|
83
|
-
```
|
|
84
|
-
|
|
85
|
-
---
|
|
86
|
-
|
|
87
|
-
## Output
|
|
88
|
-
|
|
89
|
-
The tool generates a tab-delimited output file with the following columns:
|
|
90
|
-
|
|
91
|
-
1. **Sample:** Name of the processed `.tsv` file.
|
|
92
|
-
2. **Taxa:** Genus-level taxonomic assignment extracted from the `Lineage` column.
|
|
93
|
-
3. **Proportion:** Proportion of matches to the given functional ID within the sample.
|
|
94
|
-
|
|
95
|
-
Example output:
|
|
96
|
-
|
|
97
|
-
```
|
|
98
|
-
Function ID: GO:0002
|
|
99
|
-
Sample Taxa Reads Assigned (Function) Proportion (Function) Proportion (Total Reads)
|
|
100
|
-
PN0536_0003_S83_Final_Contig.tsv Gordonibacter 60788 0.075 0.002
|
|
101
|
-
PN0536_0003_S83_Final_Contig.tsv Streptomyces 115671 0.142 0.004
|
|
102
|
-
PN0536_0003_S83_Final_Contig.tsv unknown 80890 0.099 0.003
|
|
103
|
-
PN0536_0003_S83_Final_Contig.tsv Clostridium 51018 0.063 0.002
|
|
104
|
-
PN0536_0003_S83_Final_Contig.tsv Lactobacillus 149909 0.184 0.005
|
|
105
|
-
PN0536_0003_S83_Final_Contig.tsv Limosilactobacillus 79694 0.098 0.003
|
|
106
|
-
```
|
|
107
|
-
|
|
108
|
-
---
|
|
109
|
-
|
|
110
|
-
## Implementation Details
|
|
111
|
-
|
|
112
|
-
### Workflow
|
|
113
|
-
|
|
114
|
-
1. The script reads `.tsv` files from the specified directory.
|
|
115
|
-
2. For each file, it searches for occurrences of the given functional ID within specific columns.
|
|
116
|
-
3. Matches are associated with genus-level taxonomic information extracted from the `Lineage` column.
|
|
117
|
-
4. Taxa proportions are calculated and saved to the output file.
|
|
118
|
-
|
|
119
|
-
### Large File Handling (Might be a failure point)
|
|
120
|
-
|
|
121
|
-
The script uses `csv.field_size_limit` to handle exceptionally large `.tsv` files.
|
|
122
|
-
|
|
123
|
-
---
|
|
124
|
-
|
|
125
|
-
## Future Plans
|
|
126
|
-
|
|
127
|
-
- Add support for additional file formats (e.g., `.csv`, `.txt`).
|
|
128
|
-
- Expand functionality for more complex taxonomic and functional analyses.
|
|
129
|
-
---
|
|
130
|
-
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|