MetaPont 0.0.2__tar.gz → 0.0.3__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,210 @@
1
+ Metadata-Version: 2.1
2
+ Name: MetaPont
3
+ Version: 0.0.3
4
+ Summary: MetaPont - A tool to bridge the gap between the output of metagenomic tools and the analysis of the data
5
+ Home-page: https://github.com/TheHuwsLab/MetaPont
6
+ Author: Nicholas Dimonaco
7
+ Author-email: nicholas@dimonaco.co.uk
8
+ Project-URL: Bug Tracker, https://github.com/TheHuwsLab/MetaPont/issues
9
+ Classifier: Programming Language :: Python :: 3
10
+ Classifier: License :: OSI Approved :: GNU General Public License v3 (GPLv3)
11
+ Classifier: Operating System :: OS Independent
12
+ Requires-Python: >=3.6
13
+ Description-Content-Type: text/markdown
14
+ License-File: LICENSE
15
+
16
+ # MetaPont
17
+ **MetaPont** - A tool to bridge the gap between the output of metagenomic tools and the analysis of the data
18
+
19
+ ## Features - These are the current aims of this project - Still under development
20
+
21
+ - **Targeted Functional Analysis:** Search for specific functional IDs (e.g., GO terms) within the `_Final_Contig.tsv` files provided by the HuwsLab Metagenome Workflow (https://github.com/TheHuwsLab/Metagenome_Workflow) .
22
+ - **Taxonomic Breakdown:** Extract genus-level taxonomy information and calculate their proportions in the dataset.
23
+ - **Batch Processing:** Analyse all `_Final_Contig.tsv` files in a specified directory.
24
+ - **Customisable Output:** Save results in a format suitable for downstream analysis.
25
+
26
+ ---
27
+
28
+ ## Installation
29
+
30
+ ### Prerequisites
31
+
32
+ Ensure you have the following installed:
33
+
34
+ - Python ~3.6 or later
35
+
36
+ ### Installation via pip
37
+
38
+ MetaPont is provided as a pip distribution.
39
+
40
+ ```bash
41
+ pip install MetaPont
42
+ ```
43
+
44
+ ---
45
+
46
+ ## Usage
47
+
48
+ ---
49
+ ### Extract-By-Function Command-line Arguments
50
+ ```Extract-By-Function -h ```
51
+ ```bash
52
+ usage: Extract_By_Function.py [-h] -d DIRECTORY -f FUNCTION_ID -o OUTPUT
53
+ [-m MIN_PROPORTION] [-top TOP_TAXA]
54
+
55
+ MetaPont v0.0.3: Extract-By-Function - Identify taxa contributing to a
56
+ specific function.
57
+
58
+ options:
59
+ -h, --help show this help message and exit
60
+ -d DIRECTORY, --directory DIRECTORY
61
+ Directory containing TSV files to analyse.
62
+ -f FUNCTION_ID, --function_id FUNCTION_ID
63
+ Specific function ID to search for (e.g.,
64
+ 'GO:0016597').
65
+ -o OUTPUT, --output OUTPUT
66
+ Output file to save results.
67
+ -m MIN_PROPORTION, --min_proportion MIN_PROPORTION
68
+ Minimum proportion threshold for taxa to be included
69
+ in the output.
70
+ -top TOP_TAXA, --top_taxa TOP_TAXA
71
+ Top n taxa to be included in the output.
72
+
73
+ ```
74
+
75
+ The `Extract-By-Function` tool provides several command-line options: \
76
+ Note: Either -m or -top is required.
77
+
78
+ | Option | Description | Required | Default |
79
+ |--------------------------|------------------------------------------------------------|----------|---------|
80
+ | `-d`, `--directory` | Directory containing `_Final_Contig.tsv` files to analyse. | Yes | None |
81
+ | `-f`, `--function_id` | Functional ID to search for (e.g., `GO:0016597`). | Yes | None |
82
+ | `-m`, `--min_proportion` | Minimum proportion needed for reporting. | Yes/No | None |
83
+ | `-top`, `--top_taxa` | Number of taxa to report. | Yes/No | None |
84
+ | `-o`, `--output` | Output file name to save results. | Yes | None |
85
+
86
+ ### Example
87
+
88
+ To search for the functional ID `GO:0016597` in all `_Final_Contig.tsv` files within the `test_data/` directory:
89
+
90
+ ```bash
91
+ Extract=By-Function -d .../test_data/Final_contig/ -f GO:0016597 -top 3 -o .../test_data/Final_Contig/Extract_By_Function_Out/results.tsv
92
+ ```
93
+
94
+ ---
95
+
96
+ ## Output
97
+
98
+ The tool generates a tab-delimited output file with the following columns:
99
+
100
+ 1. **Sample:** Name of the processed Sample.
101
+ 2. **Taxa:** Genus-level taxonomic assignment extracted from the `Lineage` column.
102
+ 3. **Reads Assigned (Function):** Number of reads assigned to contigs with the given functional ID.
103
+ 3. **Proportion:** Proportion of reads assigned to contigs of stated Taxa with the given functional ID within the sample.
104
+ 4. **Proportion (Total Reads):** Proportion of reads assigned to contigs of stated Taxa with the given functional ID within the total reads of the sample.
105
+
106
+ Example output:
107
+
108
+ ```
109
+ Function ID: GO:0016597
110
+ Sample Taxa Reads Assigned (Function) Proportion (Function) Proportion (Total Reads)
111
+ PN0536_0001_S1_Final_Contig.tsv Lactobacillus 111963 0.602 0.004
112
+ PN0536_0003_S83_Final_Contig.tsv Lactobacillus 20072 0.457 0.001
113
+ PN0536_0002_S2_Final_Contig.tsv Acutalibacter 145222 0.795 0.005
114
+ PN0536_0004_S3_Final_Contig.tsv Lactobacillus 40076 0.404 0.002
115
+ ```
116
+
117
+ ---
118
+
119
+
120
+
121
+ ### Workflow - unfinished
122
+
123
+ 1. The script reads `_Final_Contig.tsv` files from the specified directory.
124
+ 2. For each file, it searches for occurrences of the given functional ID within specific columns.
125
+ 3. Matches are associated with genus-level taxonomic information extracted from the `Lineage` column.
126
+ 4. Taxa proportions are calculated and saved to the output file.
127
+
128
+ ---
129
+ ## Extract-By-Taxa Command-line Arguments
130
+ ```Extract-By-Taxa -h ```
131
+ ```bash
132
+ usage: Extract_By_Taxa.py [-h] -d DIRECTORY -t TAXON -o OUTPUT -func
133
+ FUNCTIONAL_CLASSES [-top TOP_FUNCTIONS]
134
+
135
+ MetaPont: Extract Top Functions by Taxon
136
+
137
+ options:
138
+ -h, --help show this help message and exit
139
+ -d DIRECTORY, --directory DIRECTORY
140
+ Directory containing TSV files to analyse.
141
+ -t TAXON, --taxon TAXON
142
+ Target taxon to search for (e.g., 'g__Bacillus').
143
+ -o OUTPUT, --output OUTPUT
144
+ Output file to save results.
145
+ -func FUNCTIONAL_CLASSES, --functional_classes FUNCTIONAL_CLASSES
146
+ Which functional classes to report (e.g. GO,EC,KEGG
147
+ etc).
148
+ -top TOP_FUNCTIONS, --top_functions TOP_FUNCTIONS
149
+ Top n functions to include in the output for each
150
+ sample (default: 3).
151
+
152
+ ```
153
+
154
+ The `Extract-By-Taxa` tool provides several command-line options:
155
+
156
+
157
+ | Option | Description | Required | Default |
158
+ |-----------------------------------|-------------------------------------------------------------|----------|---------|
159
+ | `-d`, `--directory` | Directory containing `_Fincal_Contig.tsv` files to analyse. | Yes | None |
160
+ | `-t`, `--taxon` | Taxa to search for (e.g., `g__Bacillus`). | Yes | None |
161
+ | `-func`, `--functional_classes` | Functional classes to report (e.g. GO,EC,KEGG etc). | Yes | None |
162
+ | `-top`, `--top_taxa` | Number of functions to report (default 3). | No | None |
163
+ | `-o`, `--output` | Output file name to save results. | Yes | None |
164
+
165
+ ### Example
166
+
167
+ To search for the top reported functions for taxon `g__Bacillus` in all `_Final_Contig.tsv` files within the `test_data/` directory:
168
+
169
+ ```bash
170
+ Extract-By-Taxa -d .../test_data/Final_Contig -t g__Bacillus -o .../test_data/Final_Contig/Extract_By_Taxa/results.tsv -func GO
171
+ ```
172
+
173
+ ---
174
+ ## Output
175
+
176
+ The tool generates a tab-delimited output file with the following columns:
177
+
178
+ 1. **Sample:** Name of the processed Sample.
179
+ 2. **Function:** Reported 'top' function.
180
+ 3. **Num of Assignments (Functions):** Number of times the function has been assigned across all contigs reported as chosen Taxon.
181
+
182
+ Example output:
183
+
184
+ ```
185
+ Selected Taxon: g__Bacillus
186
+ Sample Function Num of Assignments
187
+ PN0536_0001_S1 GO:0008150 296
188
+ PN0536_0001_S1 GO:0003674 285
189
+ PN0536_0001_S1 GO:0005575 254
190
+ PN0536_0003_S83 GO:0005575 45
191
+ PN0536_0003_S83 GO:0008150 44
192
+ PN0536_0003_S83 GO:0003674 43
193
+ PN0536_0002_S2 GO:0005575 5
194
+ PN0536_0002_S2 GO:0008150 5
195
+ PN0536_0002_S2 GO:0005623 4
196
+ PN0536_0004_S3 GO:0008150 4
197
+ PN0536_0004_S3 GO:0003674 3
198
+ PN0536_0004_S3 GO:0005488 3
199
+
200
+ ```
201
+
202
+ ---
203
+
204
+ ### Large File Handling (Might be a failure point)
205
+
206
+ The script uses `csv.field_size_limit` to handle exceptionally large `.tsv` files.
207
+
208
+ ---
209
+
210
+
@@ -0,0 +1,195 @@
1
+ # MetaPont
2
+ **MetaPont** - A tool to bridge the gap between the output of metagenomic tools and the analysis of the data
3
+
4
+ ## Features - These are the current aims of this project - Still under development
5
+
6
+ - **Targeted Functional Analysis:** Search for specific functional IDs (e.g., GO terms) within the `_Final_Contig.tsv` files provided by the HuwsLab Metagenome Workflow (https://github.com/TheHuwsLab/Metagenome_Workflow) .
7
+ - **Taxonomic Breakdown:** Extract genus-level taxonomy information and calculate their proportions in the dataset.
8
+ - **Batch Processing:** Analyse all `_Final_Contig.tsv` files in a specified directory.
9
+ - **Customisable Output:** Save results in a format suitable for downstream analysis.
10
+
11
+ ---
12
+
13
+ ## Installation
14
+
15
+ ### Prerequisites
16
+
17
+ Ensure you have the following installed:
18
+
19
+ - Python ~3.6 or later
20
+
21
+ ### Installation via pip
22
+
23
+ MetaPont is provided as a pip distribution.
24
+
25
+ ```bash
26
+ pip install MetaPont
27
+ ```
28
+
29
+ ---
30
+
31
+ ## Usage
32
+
33
+ ---
34
+ ### Extract-By-Function Command-line Arguments
35
+ ```Extract-By-Function -h ```
36
+ ```bash
37
+ usage: Extract_By_Function.py [-h] -d DIRECTORY -f FUNCTION_ID -o OUTPUT
38
+ [-m MIN_PROPORTION] [-top TOP_TAXA]
39
+
40
+ MetaPont v0.0.3: Extract-By-Function - Identify taxa contributing to a
41
+ specific function.
42
+
43
+ options:
44
+ -h, --help show this help message and exit
45
+ -d DIRECTORY, --directory DIRECTORY
46
+ Directory containing TSV files to analyse.
47
+ -f FUNCTION_ID, --function_id FUNCTION_ID
48
+ Specific function ID to search for (e.g.,
49
+ 'GO:0016597').
50
+ -o OUTPUT, --output OUTPUT
51
+ Output file to save results.
52
+ -m MIN_PROPORTION, --min_proportion MIN_PROPORTION
53
+ Minimum proportion threshold for taxa to be included
54
+ in the output.
55
+ -top TOP_TAXA, --top_taxa TOP_TAXA
56
+ Top n taxa to be included in the output.
57
+
58
+ ```
59
+
60
+ The `Extract-By-Function` tool provides several command-line options: \
61
+ Note: Either -m or -top is required.
62
+
63
+ | Option | Description | Required | Default |
64
+ |--------------------------|------------------------------------------------------------|----------|---------|
65
+ | `-d`, `--directory` | Directory containing `_Final_Contig.tsv` files to analyse. | Yes | None |
66
+ | `-f`, `--function_id` | Functional ID to search for (e.g., `GO:0016597`). | Yes | None |
67
+ | `-m`, `--min_proportion` | Minimum proportion needed for reporting. | Yes/No | None |
68
+ | `-top`, `--top_taxa` | Number of taxa to report. | Yes/No | None |
69
+ | `-o`, `--output` | Output file name to save results. | Yes | None |
70
+
71
+ ### Example
72
+
73
+ To search for the functional ID `GO:0016597` in all `_Final_Contig.tsv` files within the `test_data/` directory:
74
+
75
+ ```bash
76
+ Extract=By-Function -d .../test_data/Final_contig/ -f GO:0016597 -top 3 -o .../test_data/Final_Contig/Extract_By_Function_Out/results.tsv
77
+ ```
78
+
79
+ ---
80
+
81
+ ## Output
82
+
83
+ The tool generates a tab-delimited output file with the following columns:
84
+
85
+ 1. **Sample:** Name of the processed Sample.
86
+ 2. **Taxa:** Genus-level taxonomic assignment extracted from the `Lineage` column.
87
+ 3. **Reads Assigned (Function):** Number of reads assigned to contigs with the given functional ID.
88
+ 3. **Proportion:** Proportion of reads assigned to contigs of stated Taxa with the given functional ID within the sample.
89
+ 4. **Proportion (Total Reads):** Proportion of reads assigned to contigs of stated Taxa with the given functional ID within the total reads of the sample.
90
+
91
+ Example output:
92
+
93
+ ```
94
+ Function ID: GO:0016597
95
+ Sample Taxa Reads Assigned (Function) Proportion (Function) Proportion (Total Reads)
96
+ PN0536_0001_S1_Final_Contig.tsv Lactobacillus 111963 0.602 0.004
97
+ PN0536_0003_S83_Final_Contig.tsv Lactobacillus 20072 0.457 0.001
98
+ PN0536_0002_S2_Final_Contig.tsv Acutalibacter 145222 0.795 0.005
99
+ PN0536_0004_S3_Final_Contig.tsv Lactobacillus 40076 0.404 0.002
100
+ ```
101
+
102
+ ---
103
+
104
+
105
+
106
+ ### Workflow - unfinished
107
+
108
+ 1. The script reads `_Final_Contig.tsv` files from the specified directory.
109
+ 2. For each file, it searches for occurrences of the given functional ID within specific columns.
110
+ 3. Matches are associated with genus-level taxonomic information extracted from the `Lineage` column.
111
+ 4. Taxa proportions are calculated and saved to the output file.
112
+
113
+ ---
114
+ ## Extract-By-Taxa Command-line Arguments
115
+ ```Extract-By-Taxa -h ```
116
+ ```bash
117
+ usage: Extract_By_Taxa.py [-h] -d DIRECTORY -t TAXON -o OUTPUT -func
118
+ FUNCTIONAL_CLASSES [-top TOP_FUNCTIONS]
119
+
120
+ MetaPont: Extract Top Functions by Taxon
121
+
122
+ options:
123
+ -h, --help show this help message and exit
124
+ -d DIRECTORY, --directory DIRECTORY
125
+ Directory containing TSV files to analyse.
126
+ -t TAXON, --taxon TAXON
127
+ Target taxon to search for (e.g., 'g__Bacillus').
128
+ -o OUTPUT, --output OUTPUT
129
+ Output file to save results.
130
+ -func FUNCTIONAL_CLASSES, --functional_classes FUNCTIONAL_CLASSES
131
+ Which functional classes to report (e.g. GO,EC,KEGG
132
+ etc).
133
+ -top TOP_FUNCTIONS, --top_functions TOP_FUNCTIONS
134
+ Top n functions to include in the output for each
135
+ sample (default: 3).
136
+
137
+ ```
138
+
139
+ The `Extract-By-Taxa` tool provides several command-line options:
140
+
141
+
142
+ | Option | Description | Required | Default |
143
+ |-----------------------------------|-------------------------------------------------------------|----------|---------|
144
+ | `-d`, `--directory` | Directory containing `_Fincal_Contig.tsv` files to analyse. | Yes | None |
145
+ | `-t`, `--taxon` | Taxa to search for (e.g., `g__Bacillus`). | Yes | None |
146
+ | `-func`, `--functional_classes` | Functional classes to report (e.g. GO,EC,KEGG etc). | Yes | None |
147
+ | `-top`, `--top_taxa` | Number of functions to report (default 3). | No | None |
148
+ | `-o`, `--output` | Output file name to save results. | Yes | None |
149
+
150
+ ### Example
151
+
152
+ To search for the top reported functions for taxon `g__Bacillus` in all `_Final_Contig.tsv` files within the `test_data/` directory:
153
+
154
+ ```bash
155
+ Extract-By-Taxa -d .../test_data/Final_Contig -t g__Bacillus -o .../test_data/Final_Contig/Extract_By_Taxa/results.tsv -func GO
156
+ ```
157
+
158
+ ---
159
+ ## Output
160
+
161
+ The tool generates a tab-delimited output file with the following columns:
162
+
163
+ 1. **Sample:** Name of the processed Sample.
164
+ 2. **Function:** Reported 'top' function.
165
+ 3. **Num of Assignments (Functions):** Number of times the function has been assigned across all contigs reported as chosen Taxon.
166
+
167
+ Example output:
168
+
169
+ ```
170
+ Selected Taxon: g__Bacillus
171
+ Sample Function Num of Assignments
172
+ PN0536_0001_S1 GO:0008150 296
173
+ PN0536_0001_S1 GO:0003674 285
174
+ PN0536_0001_S1 GO:0005575 254
175
+ PN0536_0003_S83 GO:0005575 45
176
+ PN0536_0003_S83 GO:0008150 44
177
+ PN0536_0003_S83 GO:0003674 43
178
+ PN0536_0002_S2 GO:0005575 5
179
+ PN0536_0002_S2 GO:0008150 5
180
+ PN0536_0002_S2 GO:0005623 4
181
+ PN0536_0004_S3 GO:0008150 4
182
+ PN0536_0004_S3 GO:0003674 3
183
+ PN0536_0004_S3 GO:0005488 3
184
+
185
+ ```
186
+
187
+ ---
188
+
189
+ ### Large File Handling (Might be a failure point)
190
+
191
+ The script uses `csv.field_size_limit` to handle exceptionally large `.tsv` files.
192
+
193
+ ---
194
+
195
+
@@ -1,6 +1,6 @@
1
1
  [metadata]
2
2
  name = MetaPont
3
- version = v0.0.2
3
+ version = v0.0.3
4
4
  author = Nicholas Dimonaco
5
5
  author_email = nicholas@dimonaco.co.uk
6
6
  description = MetaPont - A tool to bridge the gap between the output of metagenomic tools and the analysis of the data
@@ -28,6 +28,7 @@ include = *
28
28
  [options.entry_points]
29
29
  console_scripts =
30
30
  Extract-By-Function = MetaPont.Extract_By_Function:main
31
+ Extract-By-Taxa = MetaPont.Extract_By_Taxa:main
31
32
 
32
33
  [egg_info]
33
34
  tag_build =
@@ -62,18 +62,26 @@ def main():
62
62
  )
63
63
  parser.add_argument(
64
64
  "-f", "--function_id", required=True,
65
- help="Specific function ID to search for (e.g., 'GO:0002')."
65
+ help="Specific function ID to search for (e.g., 'GO:0016597')."
66
66
  )
67
67
  parser.add_argument(
68
- "-o", "--output", default="output_taxa_details.tsv",
69
- help="Output file to save results (default: output_taxa_details.tsv)."
68
+ "-o", "--output", required=True,
69
+ help="Output file to save results."
70
70
  )
71
71
  parser.add_argument(
72
- "-m", "--min_proportion", type=float, default=0.05,
73
- help="Minimum proportion threshold for taxa to be included in the output (default: 0.05)."
72
+ "-m", "--min_proportion", type=float,
73
+ help="Minimum proportion threshold for taxa to be included in the output."
74
+ )
75
+ parser.add_argument(
76
+ "-top", "--top_taxa", type=int,
77
+ help="Top n taxa to be included in the output."
74
78
  )
75
79
 
76
80
  options = parser.parse_args()
81
+ if not options.min_proportion and not options.top_taxa:
82
+ sys.exit("Error: Please specify either a minimum proportion or the number of top taxa to include in the output.")
83
+ elif options.min_proportion and options.top_taxa:
84
+ sys.exit("Error: Please specify either a minimum proportion or the number of top taxa to include in the output, not both.")
77
85
  print("Running MetaPont: Extract-By-Function " + MetaPont_Version)
78
86
 
79
87
  input_path = os.path.abspath(options.directory)
@@ -94,12 +102,21 @@ def main():
94
102
  out.write("Function ID: " + options.function_id + "\n")
95
103
  out.write("Sample\tTaxa\tReads Assigned (Function)\tProportion (Function)\tProportion (Total Reads)\n")
96
104
  for sample, (taxa_reads_function, total_reads_function, total_reads_all) in all_results.items():
97
- for taxa, reads_function in taxa_reads_function.items():
98
- proportion_function = reads_function / total_reads_function if total_reads_function > 0 else 0
99
- proportion_total = reads_function / total_reads_all if total_reads_all > 0 else 0
100
-
101
- if proportion_function >= options.min_proportion: # Apply minimum proportion filter
102
- out.write(f"{sample}\t{taxa}\t{reads_function}\t{proportion_function:.3f}\t{proportion_total:.3f}\n")
105
+ if options.min_proportion:
106
+ for taxa, reads_function in taxa_reads_function.items():
107
+ proportion_function = reads_function / total_reads_function if total_reads_function > 0 else 0
108
+ proportion_total = reads_function / total_reads_all if total_reads_all > 0 else 0
109
+ if proportion_function >= options.min_proportion: # Apply minimum proportion filter
110
+ out.write(f"{sample}\t{taxa}\t{reads_function}\t{proportion_function:.3f}\t{proportion_total:.3f}\n")
111
+ elif options.top_taxa:
112
+ sorted_taxa_reads = sorted(taxa_reads_function.items(), key=lambda x: x[1], reverse=True)
113
+ for i, (taxa, reads_function) in enumerate(sorted_taxa_reads):
114
+ if i < options.top_taxa:
115
+ proportion_function = reads_function / total_reads_function if total_reads_function > 0 else 0
116
+ proportion_total = reads_function / total_reads_all if total_reads_all > 0 else 0
117
+ out.write(f"{sample.replace('_Final_Contig.tsv','')}\t{taxa}\t{reads_function}\t{proportion_function:.3f}\t{proportion_total:.3f}\n")
118
+ else:
119
+ break
103
120
 
104
121
  print(f"Results saved to {options.output}")
105
122
 
@@ -0,0 +1,130 @@
1
+ import argparse
2
+ import os
3
+ import csv
4
+ import sys
5
+ from collections import defaultdict
6
+
7
+ # Adjust constants import for module vs. standalone script usage
8
+ try:
9
+ from .constants import *
10
+ except (ModuleNotFoundError, ImportError, NameError, TypeError):
11
+ from constants import *
12
+
13
+ # Needed to account for the large CSV/TSV files
14
+ csv.field_size_limit(sys.maxsize)
15
+
16
+
17
+ def process_tsv_by_taxon(file_path, target_taxon):
18
+ """
19
+ Processes a TSV file to calculate the top functions grouped by taxon.
20
+
21
+ Parameters:
22
+ file_path (str): Path to the TSV file.
23
+ target_taxon (str): Taxon to search for in the lineage column.
24
+
25
+ Returns:
26
+ taxon_function_reads (dict): A dictionary mapping functions to their total reads for the specified taxon.
27
+ total_reads_taxon (int): Total reads assigned to the specified taxon across all functions.
28
+ """
29
+ taxon_function_presence = defaultdict(int) # Reads for each function for the specified taxon
30
+ total_reads_taxon = 0 # Total reads for the specified taxon
31
+
32
+ with open(file_path, "r") as tsv_file:
33
+ reader = csv.reader(tsv_file, delimiter="\t")
34
+ next(reader) # Skip the first row (sample name)
35
+ headers = next(reader) # Read headers from the second row
36
+
37
+ # Locate the necessary columns
38
+ taxa_idx = headers.index("Lineage")
39
+ reads_idx = headers.index("Mapped_Reads")
40
+
41
+ # Process each row
42
+ for idx, row in enumerate(reader):
43
+ if len(row) < len(headers):
44
+ continue # Skip malformed rows
45
+
46
+ lineage = row[taxa_idx]
47
+ reads = int(row[reads_idx]) # Get the number of reads for this row
48
+
49
+ # Check if this row matches the specified taxon
50
+ if target_taxon in lineage:
51
+ total_reads_taxon += reads # Increment total reads for the taxon
52
+
53
+ # Loop through the function columns (row[6:] onward)
54
+ for col_idx, cell in enumerate(row[6:], start=6):
55
+ if cell:
56
+ header = headers[col_idx]
57
+ if header not in taxon_function_presence:
58
+ taxon_function_presence[header] = defaultdict(int)
59
+ for function in cell.replace(',', '|').split('|'):
60
+ taxon_function_presence[header][function] += 1
61
+
62
+ return taxon_function_presence, total_reads_taxon
63
+
64
+
65
+ def main():
66
+ parser = argparse.ArgumentParser(description='MetaPont: Extract Top Functions by Taxon')
67
+ parser.add_argument(
68
+ "-d", "--directory", required=True,
69
+ help="Directory containing TSV files to analyse."
70
+ )
71
+ parser.add_argument(
72
+ "-t", "--taxon", required=True,
73
+ help="Target taxon to search for (e.g., 'g__Escherichia')."
74
+ )
75
+ parser.add_argument(
76
+ "-o", "--output", required=True,
77
+ help="Output file to save results."
78
+ )
79
+ parser.add_argument(
80
+ "-func", "--functional_classes", required=True,
81
+ help="Which functional classes to report (e.g. GO,EC,KEGG etc)."
82
+ )
83
+ parser.add_argument(
84
+ "-top", "--top_functions", type=int, default=3,
85
+ help="Top n functions to include in the output for each sample (default: 3)."
86
+ )
87
+
88
+ options = parser.parse_args()
89
+ print("Running MetaPont: Extract Top Functions by Taxon")
90
+
91
+ input_path = os.path.abspath(options.directory)
92
+ output_path = os.path.abspath(options.output)
93
+
94
+ all_results = {}
95
+
96
+ # Process each TSV file in the directory
97
+ for file_name in os.listdir(input_path):
98
+ if file_name.endswith("_Final_Contig.tsv"):
99
+ file_path = os.path.join(options.directory, file_name)
100
+ print(f"Processing file: {file_name}")
101
+ taxon_function_presence, total_reads_taxon = process_tsv_by_taxon(file_path, options.taxon)
102
+ all_results[file_name] = (taxon_function_presence, total_reads_taxon)
103
+
104
+ # Write results to output
105
+ with open(output_path, "w") as out:
106
+ out.write("Selected Taxon: " + options.taxon + "\n")
107
+ out.write("Sample\tFunction\tNum of Assignments\n")
108
+ for sample, (taxon_function_presence, total_assignments_taxon) in all_results.items():
109
+ for function, assignments_dict in taxon_function_presence.items():
110
+ if any(func_class in function for func_class in options.functional_classes.split(',')):
111
+ sorted_functions = sorted(assignments_dict.items(), key=lambda x: x[1], reverse=True)
112
+ for i, (sub_function, assignments) in enumerate(sorted_functions):
113
+ if i < options.top_functions:
114
+ #proportion_taxon = assignments / total_reads_taxon if total_reads_taxon > 0 else 0
115
+ out.write(f"{sample.replace('_Final_Contig.tsv', '')}\t{sub_function}\t{assignments}\n")
116
+ else:
117
+ break
118
+ # sorted_functions = sorted(taxon_function_presence.items(), key=lambda x: x[1], reverse=True)
119
+ # for i, (function, reads) in enumerate(sorted_functions):
120
+ # if i < options.top_functions:
121
+ # proportion_taxon = reads / total_reads_taxon if total_reads_taxon > 0 else 0
122
+ # out.write(f"{sample.replace('_Final_Contig.tsv', '')}\t{function}\t{reads}\t{proportion_taxon:.3f}\n")
123
+ # else:
124
+ # break
125
+
126
+ print(f"Results saved to {options.output}")
127
+
128
+
129
+ if __name__ == "__main__":
130
+ main()
@@ -0,0 +1,2 @@
1
+ MetaPont_Version = 'v0.0.3'
2
+
@@ -0,0 +1,210 @@
1
+ Metadata-Version: 2.1
2
+ Name: MetaPont
3
+ Version: 0.0.3
4
+ Summary: MetaPont - A tool to bridge the gap between the output of metagenomic tools and the analysis of the data
5
+ Home-page: https://github.com/TheHuwsLab/MetaPont
6
+ Author: Nicholas Dimonaco
7
+ Author-email: nicholas@dimonaco.co.uk
8
+ Project-URL: Bug Tracker, https://github.com/TheHuwsLab/MetaPont/issues
9
+ Classifier: Programming Language :: Python :: 3
10
+ Classifier: License :: OSI Approved :: GNU General Public License v3 (GPLv3)
11
+ Classifier: Operating System :: OS Independent
12
+ Requires-Python: >=3.6
13
+ Description-Content-Type: text/markdown
14
+ License-File: LICENSE
15
+
16
+ # MetaPont
17
+ **MetaPont** - A tool to bridge the gap between the output of metagenomic tools and the analysis of the data
18
+
19
+ ## Features - These are the current aims of this project - Still under development
20
+
21
+ - **Targeted Functional Analysis:** Search for specific functional IDs (e.g., GO terms) within the `_Final_Contig.tsv` files provided by the HuwsLab Metagenome Workflow (https://github.com/TheHuwsLab/Metagenome_Workflow) .
22
+ - **Taxonomic Breakdown:** Extract genus-level taxonomy information and calculate their proportions in the dataset.
23
+ - **Batch Processing:** Analyse all `_Final_Contig.tsv` files in a specified directory.
24
+ - **Customisable Output:** Save results in a format suitable for downstream analysis.
25
+
26
+ ---
27
+
28
+ ## Installation
29
+
30
+ ### Prerequisites
31
+
32
+ Ensure you have the following installed:
33
+
34
+ - Python ~3.6 or later
35
+
36
+ ### Installation via pip
37
+
38
+ MetaPont is provided as a pip distribution.
39
+
40
+ ```bash
41
+ pip install MetaPont
42
+ ```
43
+
44
+ ---
45
+
46
+ ## Usage
47
+
48
+ ---
49
+ ### Extract-By-Function Command-line Arguments
50
+ ```Extract-By-Function -h ```
51
+ ```bash
52
+ usage: Extract_By_Function.py [-h] -d DIRECTORY -f FUNCTION_ID -o OUTPUT
53
+ [-m MIN_PROPORTION] [-top TOP_TAXA]
54
+
55
+ MetaPont v0.0.3: Extract-By-Function - Identify taxa contributing to a
56
+ specific function.
57
+
58
+ options:
59
+ -h, --help show this help message and exit
60
+ -d DIRECTORY, --directory DIRECTORY
61
+ Directory containing TSV files to analyse.
62
+ -f FUNCTION_ID, --function_id FUNCTION_ID
63
+ Specific function ID to search for (e.g.,
64
+ 'GO:0016597').
65
+ -o OUTPUT, --output OUTPUT
66
+ Output file to save results.
67
+ -m MIN_PROPORTION, --min_proportion MIN_PROPORTION
68
+ Minimum proportion threshold for taxa to be included
69
+ in the output.
70
+ -top TOP_TAXA, --top_taxa TOP_TAXA
71
+ Top n taxa to be included in the output.
72
+
73
+ ```
74
+
75
+ The `Extract-By-Function` tool provides several command-line options: \
76
+ Note: Either -m or -top is required.
77
+
78
+ | Option | Description | Required | Default |
79
+ |--------------------------|------------------------------------------------------------|----------|---------|
80
+ | `-d`, `--directory` | Directory containing `_Final_Contig.tsv` files to analyse. | Yes | None |
81
+ | `-f`, `--function_id` | Functional ID to search for (e.g., `GO:0016597`). | Yes | None |
82
+ | `-m`, `--min_proportion` | Minimum proportion needed for reporting. | Yes/No | None |
83
+ | `-top`, `--top_taxa` | Number of taxa to report. | Yes/No | None |
84
+ | `-o`, `--output` | Output file name to save results. | Yes | None |
85
+
86
+ ### Example
87
+
88
+ To search for the functional ID `GO:0016597` in all `_Final_Contig.tsv` files within the `test_data/` directory:
89
+
90
+ ```bash
91
+ Extract=By-Function -d .../test_data/Final_contig/ -f GO:0016597 -top 3 -o .../test_data/Final_Contig/Extract_By_Function_Out/results.tsv
92
+ ```
93
+
94
+ ---
95
+
96
+ ## Output
97
+
98
+ The tool generates a tab-delimited output file with the following columns:
99
+
100
+ 1. **Sample:** Name of the processed Sample.
101
+ 2. **Taxa:** Genus-level taxonomic assignment extracted from the `Lineage` column.
102
+ 3. **Reads Assigned (Function):** Number of reads assigned to contigs with the given functional ID.
103
+ 3. **Proportion:** Proportion of reads assigned to contigs of stated Taxa with the given functional ID within the sample.
104
+ 4. **Proportion (Total Reads):** Proportion of reads assigned to contigs of stated Taxa with the given functional ID within the total reads of the sample.
105
+
106
+ Example output:
107
+
108
+ ```
109
+ Function ID: GO:0016597
110
+ Sample Taxa Reads Assigned (Function) Proportion (Function) Proportion (Total Reads)
111
+ PN0536_0001_S1_Final_Contig.tsv Lactobacillus 111963 0.602 0.004
112
+ PN0536_0003_S83_Final_Contig.tsv Lactobacillus 20072 0.457 0.001
113
+ PN0536_0002_S2_Final_Contig.tsv Acutalibacter 145222 0.795 0.005
114
+ PN0536_0004_S3_Final_Contig.tsv Lactobacillus 40076 0.404 0.002
115
+ ```
116
+
117
+ ---
118
+
119
+
120
+
121
+ ### Workflow - unfinished
122
+
123
+ 1. The script reads `_Final_Contig.tsv` files from the specified directory.
124
+ 2. For each file, it searches for occurrences of the given functional ID within specific columns.
125
+ 3. Matches are associated with genus-level taxonomic information extracted from the `Lineage` column.
126
+ 4. Taxa proportions are calculated and saved to the output file.
127
+
128
+ ---
129
+ ## Extract-By-Taxa Command-line Arguments
130
+ ```Extract-By-Taxa -h ```
131
+ ```bash
132
+ usage: Extract_By_Taxa.py [-h] -d DIRECTORY -t TAXON -o OUTPUT -func
133
+ FUNCTIONAL_CLASSES [-top TOP_FUNCTIONS]
134
+
135
+ MetaPont: Extract Top Functions by Taxon
136
+
137
+ options:
138
+ -h, --help show this help message and exit
139
+ -d DIRECTORY, --directory DIRECTORY
140
+ Directory containing TSV files to analyse.
141
+ -t TAXON, --taxon TAXON
142
+ Target taxon to search for (e.g., 'g__Bacillus').
143
+ -o OUTPUT, --output OUTPUT
144
+ Output file to save results.
145
+ -func FUNCTIONAL_CLASSES, --functional_classes FUNCTIONAL_CLASSES
146
+ Which functional classes to report (e.g. GO,EC,KEGG
147
+ etc).
148
+ -top TOP_FUNCTIONS, --top_functions TOP_FUNCTIONS
149
+ Top n functions to include in the output for each
150
+ sample (default: 3).
151
+
152
+ ```
153
+
154
+ The `Extract-By-Taxa` tool provides several command-line options:
155
+
156
+
157
+ | Option | Description | Required | Default |
158
+ |-----------------------------------|-------------------------------------------------------------|----------|---------|
159
+ | `-d`, `--directory` | Directory containing `_Fincal_Contig.tsv` files to analyse. | Yes | None |
160
+ | `-t`, `--taxon` | Taxa to search for (e.g., `g__Bacillus`). | Yes | None |
161
+ | `-func`, `--functional_classes` | Functional classes to report (e.g. GO,EC,KEGG etc). | Yes | None |
162
+ | `-top`, `--top_taxa` | Number of functions to report (default 3). | No | None |
163
+ | `-o`, `--output` | Output file name to save results. | Yes | None |
164
+
165
+ ### Example
166
+
167
+ To search for the top reported functions for taxon `g__Bacillus` in all `_Final_Contig.tsv` files within the `test_data/` directory:
168
+
169
+ ```bash
170
+ Extract-By-Taxa -d .../test_data/Final_Contig -t g__Bacillus -o .../test_data/Final_Contig/Extract_By_Taxa/results.tsv -func GO
171
+ ```
172
+
173
+ ---
174
+ ## Output
175
+
176
+ The tool generates a tab-delimited output file with the following columns:
177
+
178
+ 1. **Sample:** Name of the processed Sample.
179
+ 2. **Function:** Reported 'top' function.
180
+ 3. **Num of Assignments (Functions):** Number of times the function has been assigned across all contigs reported as chosen Taxon.
181
+
182
+ Example output:
183
+
184
+ ```
185
+ Selected Taxon: g__Bacillus
186
+ Sample Function Num of Assignments
187
+ PN0536_0001_S1 GO:0008150 296
188
+ PN0536_0001_S1 GO:0003674 285
189
+ PN0536_0001_S1 GO:0005575 254
190
+ PN0536_0003_S83 GO:0005575 45
191
+ PN0536_0003_S83 GO:0008150 44
192
+ PN0536_0003_S83 GO:0003674 43
193
+ PN0536_0002_S2 GO:0005575 5
194
+ PN0536_0002_S2 GO:0008150 5
195
+ PN0536_0002_S2 GO:0005623 4
196
+ PN0536_0004_S3 GO:0008150 4
197
+ PN0536_0004_S3 GO:0003674 3
198
+ PN0536_0004_S3 GO:0005488 3
199
+
200
+ ```
201
+
202
+ ---
203
+
204
+ ### Large File Handling (Might be a failure point)
205
+
206
+ The script uses `csv.field_size_limit` to handle exceptionally large `.tsv` files.
207
+
208
+ ---
209
+
210
+
@@ -3,6 +3,7 @@ README.md
3
3
  pyproject.toml
4
4
  setup.cfg
5
5
  src/MetaPont/Extract_By_Function.py
6
+ src/MetaPont/Extract_By_Taxa.py
6
7
  src/MetaPont/Function_By_Taxa.py
7
8
  src/MetaPont/__init__.py
8
9
  src/MetaPont/constants.py
@@ -1,2 +1,3 @@
1
1
  [console_scripts]
2
2
  Extract-By-Function = MetaPont.Extract_By_Function:main
3
+ Extract-By-Taxa = MetaPont.Extract_By_Taxa:main
metapont-0.0.2/PKG-INFO DELETED
@@ -1,130 +0,0 @@
1
- Metadata-Version: 2.1
2
- Name: MetaPont
3
- Version: 0.0.2
4
- Summary: MetaPont - A tool to bridge the gap between the output of metagenomic tools and the analysis of the data
5
- Home-page: https://github.com/TheHuwsLab/MetaPont
6
- Author: Nicholas Dimonaco
7
- Author-email: nicholas@dimonaco.co.uk
8
- Project-URL: Bug Tracker, https://github.com/TheHuwsLab/MetaPont/issues
9
- Classifier: Programming Language :: Python :: 3
10
- Classifier: License :: OSI Approved :: GNU General Public License v3 (GPLv3)
11
- Classifier: Operating System :: OS Independent
12
- Requires-Python: >=3.6
13
- Description-Content-Type: text/markdown
14
- License-File: LICENSE
15
-
16
- # MetaPont
17
- **MetaPont** - A tool to bridge the gap between the output of metagenomic tools and the analysis of the data
18
-
19
- ## Features - These are the current aims of this project - Still under development
20
-
21
- - **Targeted Functional Analysis:** Search for specific functional IDs (e.g., GO terms) within the `.tsv` files provided by the HuwsLab Metagenome Workflow (https://github.com/TheHuwsLab/Metagenome_Workflow) .
22
- - **Taxonomic Breakdown:** Extract genus-level taxonomy information and calculate their proportions in the dataset.
23
- - **Batch Processing:** Analyse all `.tsv` files in a specified directory.
24
- - **Customisable Output:** Save results in a format suitable for downstream analysis.
25
-
26
- ---
27
-
28
- ## Installation
29
-
30
- ### Prerequisites
31
-
32
- Ensure you have the following installed:
33
-
34
- - Python ~3.6 or later
35
- - Required Python libraries: `argparse`, `csv`, and `collections` (standard libs).
36
-
37
- ### Installation via pip
38
-
39
- MetaPont is provided as a pip distribution.
40
-
41
- ```bash
42
- pip install MetaPont
43
- ```
44
-
45
- ---
46
-
47
- ## Usage
48
-
49
- ### Command-line Arguments
50
- ```Extract-By-Function -h ```
51
- ```bash
52
- usage: Extract-By-Function [-h] -d DIRECTORY -f FUNCTION_ID [-o OUTPUT] [-m MIN_PROPORTION]
53
-
54
- MetaPont v0.0.2: Extract-By-Function - Identify taxa contributing to a specific function.
55
-
56
- options:
57
- -h, --help show this help message and exit
58
- -d DIRECTORY, --directory DIRECTORY
59
- Directory containing TSV files to analyse.
60
- -f FUNCTION_ID, --function_id FUNCTION_ID
61
- Specific function ID to search for (e.g., 'GO:0002').
62
- -o OUTPUT, --output OUTPUT
63
- Output file to save results (default: output_taxa_details.tsv).
64
- -m MIN_PROPORTION, --min_proportion MIN_PROPORTION
65
- Minimum proportion threshold for taxa to be included in the output (default: 0.05).
66
- ```
67
-
68
- The `Extract-By-Function` tool provides several command-line options:
69
-
70
- | Option | Description | Required | Default |
71
- |--------------------------|-----------------------------------------------|----------|-------------------------------|
72
- | `-d`, `--directory` | Directory containing `.tsv` files to analyse. | Yes | None |
73
- | `-f`, `--function_id` | Functional ID to search for (e.g., `GO:0002`). | Yes | None |
74
- | `-m`, `--min_proportion` | Minimum proportion needed for reporting. | Yes | 0.05 (5%) |
75
- | `-o`, `--output` | Output file name to save results. | No | `output_taxa_proportions.tsv` |
76
-
77
- ### Example
78
-
79
- To search for the functional ID `GO:0002` in all `.tsv` files within the `data/` directory:
80
-
81
- ```bash
82
- ExtractByFunction -d .../test_data/Final_contig/ -f GO:0002 -m 0.10 -o .../test_data/Final_Contig/Extract_By_Function_Out/results.tsv
83
- ```
84
-
85
- ---
86
-
87
- ## Output
88
-
89
- The tool generates a tab-delimited output file with the following columns:
90
-
91
- 1. **Sample:** Name of the processed `.tsv` file.
92
- 2. **Taxa:** Genus-level taxonomic assignment extracted from the `Lineage` column.
93
- 3. **Proportion:** Proportion of matches to the given functional ID within the sample.
94
-
95
- Example output:
96
-
97
- ```
98
- Function ID: GO:0002
99
- Sample Taxa Reads Assigned (Function) Proportion (Function) Proportion (Total Reads)
100
- PN0536_0003_S83_Final_Contig.tsv Gordonibacter 60788 0.075 0.002
101
- PN0536_0003_S83_Final_Contig.tsv Streptomyces 115671 0.142 0.004
102
- PN0536_0003_S83_Final_Contig.tsv unknown 80890 0.099 0.003
103
- PN0536_0003_S83_Final_Contig.tsv Clostridium 51018 0.063 0.002
104
- PN0536_0003_S83_Final_Contig.tsv Lactobacillus 149909 0.184 0.005
105
- PN0536_0003_S83_Final_Contig.tsv Limosilactobacillus 79694 0.098 0.003
106
- ```
107
-
108
- ---
109
-
110
- ## Implementation Details
111
-
112
- ### Workflow
113
-
114
- 1. The script reads `.tsv` files from the specified directory.
115
- 2. For each file, it searches for occurrences of the given functional ID within specific columns.
116
- 3. Matches are associated with genus-level taxonomic information extracted from the `Lineage` column.
117
- 4. Taxa proportions are calculated and saved to the output file.
118
-
119
- ### Large File Handling (Might be a failure point)
120
-
121
- The script uses `csv.field_size_limit` to handle exceptionally large `.tsv` files.
122
-
123
- ---
124
-
125
- ## Future Plans
126
-
127
- - Add support for additional file formats (e.g., `.csv`, `.txt`).
128
- - Expand functionality for more complex taxonomic and functional analyses.
129
- ---
130
-
metapont-0.0.2/README.md DELETED
@@ -1,115 +0,0 @@
1
- # MetaPont
2
- **MetaPont** - A tool to bridge the gap between the output of metagenomic tools and the analysis of the data
3
-
4
- ## Features - These are the current aims of this project - Still under development
5
-
6
- - **Targeted Functional Analysis:** Search for specific functional IDs (e.g., GO terms) within the `.tsv` files provided by the HuwsLab Metagenome Workflow (https://github.com/TheHuwsLab/Metagenome_Workflow) .
7
- - **Taxonomic Breakdown:** Extract genus-level taxonomy information and calculate their proportions in the dataset.
8
- - **Batch Processing:** Analyse all `.tsv` files in a specified directory.
9
- - **Customisable Output:** Save results in a format suitable for downstream analysis.
10
-
11
- ---
12
-
13
- ## Installation
14
-
15
- ### Prerequisites
16
-
17
- Ensure you have the following installed:
18
-
19
- - Python ~3.6 or later
20
- - Required Python libraries: `argparse`, `csv`, and `collections` (standard libs).
21
-
22
- ### Installation via pip
23
-
24
- MetaPont is provided as a pip distribution.
25
-
26
- ```bash
27
- pip install MetaPont
28
- ```
29
-
30
- ---
31
-
32
- ## Usage
33
-
34
- ### Command-line Arguments
35
- ```Extract-By-Function -h ```
36
- ```bash
37
- usage: Extract-By-Function [-h] -d DIRECTORY -f FUNCTION_ID [-o OUTPUT] [-m MIN_PROPORTION]
38
-
39
- MetaPont v0.0.2: Extract-By-Function - Identify taxa contributing to a specific function.
40
-
41
- options:
42
- -h, --help show this help message and exit
43
- -d DIRECTORY, --directory DIRECTORY
44
- Directory containing TSV files to analyse.
45
- -f FUNCTION_ID, --function_id FUNCTION_ID
46
- Specific function ID to search for (e.g., 'GO:0002').
47
- -o OUTPUT, --output OUTPUT
48
- Output file to save results (default: output_taxa_details.tsv).
49
- -m MIN_PROPORTION, --min_proportion MIN_PROPORTION
50
- Minimum proportion threshold for taxa to be included in the output (default: 0.05).
51
- ```
52
-
53
- The `Extract-By-Function` tool provides several command-line options:
54
-
55
- | Option | Description | Required | Default |
56
- |--------------------------|-----------------------------------------------|----------|-------------------------------|
57
- | `-d`, `--directory` | Directory containing `.tsv` files to analyse. | Yes | None |
58
- | `-f`, `--function_id` | Functional ID to search for (e.g., `GO:0002`). | Yes | None |
59
- | `-m`, `--min_proportion` | Minimum proportion needed for reporting. | Yes | 0.05 (5%) |
60
- | `-o`, `--output` | Output file name to save results. | No | `output_taxa_proportions.tsv` |
61
-
62
- ### Example
63
-
64
- To search for the functional ID `GO:0002` in all `.tsv` files within the `data/` directory:
65
-
66
- ```bash
67
- ExtractByFunction -d .../test_data/Final_contig/ -f GO:0002 -m 0.10 -o .../test_data/Final_Contig/Extract_By_Function_Out/results.tsv
68
- ```
69
-
70
- ---
71
-
72
- ## Output
73
-
74
- The tool generates a tab-delimited output file with the following columns:
75
-
76
- 1. **Sample:** Name of the processed `.tsv` file.
77
- 2. **Taxa:** Genus-level taxonomic assignment extracted from the `Lineage` column.
78
- 3. **Proportion:** Proportion of matches to the given functional ID within the sample.
79
-
80
- Example output:
81
-
82
- ```
83
- Function ID: GO:0002
84
- Sample Taxa Reads Assigned (Function) Proportion (Function) Proportion (Total Reads)
85
- PN0536_0003_S83_Final_Contig.tsv Gordonibacter 60788 0.075 0.002
86
- PN0536_0003_S83_Final_Contig.tsv Streptomyces 115671 0.142 0.004
87
- PN0536_0003_S83_Final_Contig.tsv unknown 80890 0.099 0.003
88
- PN0536_0003_S83_Final_Contig.tsv Clostridium 51018 0.063 0.002
89
- PN0536_0003_S83_Final_Contig.tsv Lactobacillus 149909 0.184 0.005
90
- PN0536_0003_S83_Final_Contig.tsv Limosilactobacillus 79694 0.098 0.003
91
- ```
92
-
93
- ---
94
-
95
- ## Implementation Details
96
-
97
- ### Workflow
98
-
99
- 1. The script reads `.tsv` files from the specified directory.
100
- 2. For each file, it searches for occurrences of the given functional ID within specific columns.
101
- 3. Matches are associated with genus-level taxonomic information extracted from the `Lineage` column.
102
- 4. Taxa proportions are calculated and saved to the output file.
103
-
104
- ### Large File Handling (Might be a failure point)
105
-
106
- The script uses `csv.field_size_limit` to handle exceptionally large `.tsv` files.
107
-
108
- ---
109
-
110
- ## Future Plans
111
-
112
- - Add support for additional file formats (e.g., `.csv`, `.txt`).
113
- - Expand functionality for more complex taxonomic and functional analyses.
114
- ---
115
-
@@ -1,2 +0,0 @@
1
- MetaPont_Version = 'v0.0.2'
2
-
@@ -1,130 +0,0 @@
1
- Metadata-Version: 2.1
2
- Name: MetaPont
3
- Version: 0.0.2
4
- Summary: MetaPont - A tool to bridge the gap between the output of metagenomic tools and the analysis of the data
5
- Home-page: https://github.com/TheHuwsLab/MetaPont
6
- Author: Nicholas Dimonaco
7
- Author-email: nicholas@dimonaco.co.uk
8
- Project-URL: Bug Tracker, https://github.com/TheHuwsLab/MetaPont/issues
9
- Classifier: Programming Language :: Python :: 3
10
- Classifier: License :: OSI Approved :: GNU General Public License v3 (GPLv3)
11
- Classifier: Operating System :: OS Independent
12
- Requires-Python: >=3.6
13
- Description-Content-Type: text/markdown
14
- License-File: LICENSE
15
-
16
- # MetaPont
17
- **MetaPont** - A tool to bridge the gap between the output of metagenomic tools and the analysis of the data
18
-
19
- ## Features - These are the current aims of this project - Still under development
20
-
21
- - **Targeted Functional Analysis:** Search for specific functional IDs (e.g., GO terms) within the `.tsv` files provided by the HuwsLab Metagenome Workflow (https://github.com/TheHuwsLab/Metagenome_Workflow) .
22
- - **Taxonomic Breakdown:** Extract genus-level taxonomy information and calculate their proportions in the dataset.
23
- - **Batch Processing:** Analyse all `.tsv` files in a specified directory.
24
- - **Customisable Output:** Save results in a format suitable for downstream analysis.
25
-
26
- ---
27
-
28
- ## Installation
29
-
30
- ### Prerequisites
31
-
32
- Ensure you have the following installed:
33
-
34
- - Python ~3.6 or later
35
- - Required Python libraries: `argparse`, `csv`, and `collections` (standard libs).
36
-
37
- ### Installation via pip
38
-
39
- MetaPont is provided as a pip distribution.
40
-
41
- ```bash
42
- pip install MetaPont
43
- ```
44
-
45
- ---
46
-
47
- ## Usage
48
-
49
- ### Command-line Arguments
50
- ```Extract-By-Function -h ```
51
- ```bash
52
- usage: Extract-By-Function [-h] -d DIRECTORY -f FUNCTION_ID [-o OUTPUT] [-m MIN_PROPORTION]
53
-
54
- MetaPont v0.0.2: Extract-By-Function - Identify taxa contributing to a specific function.
55
-
56
- options:
57
- -h, --help show this help message and exit
58
- -d DIRECTORY, --directory DIRECTORY
59
- Directory containing TSV files to analyse.
60
- -f FUNCTION_ID, --function_id FUNCTION_ID
61
- Specific function ID to search for (e.g., 'GO:0002').
62
- -o OUTPUT, --output OUTPUT
63
- Output file to save results (default: output_taxa_details.tsv).
64
- -m MIN_PROPORTION, --min_proportion MIN_PROPORTION
65
- Minimum proportion threshold for taxa to be included in the output (default: 0.05).
66
- ```
67
-
68
- The `Extract-By-Function` tool provides several command-line options:
69
-
70
- | Option | Description | Required | Default |
71
- |--------------------------|-----------------------------------------------|----------|-------------------------------|
72
- | `-d`, `--directory` | Directory containing `.tsv` files to analyse. | Yes | None |
73
- | `-f`, `--function_id` | Functional ID to search for (e.g., `GO:0002`). | Yes | None |
74
- | `-m`, `--min_proportion` | Minimum proportion needed for reporting. | Yes | 0.05 (5%) |
75
- | `-o`, `--output` | Output file name to save results. | No | `output_taxa_proportions.tsv` |
76
-
77
- ### Example
78
-
79
- To search for the functional ID `GO:0002` in all `.tsv` files within the `data/` directory:
80
-
81
- ```bash
82
- ExtractByFunction -d .../test_data/Final_contig/ -f GO:0002 -m 0.10 -o .../test_data/Final_Contig/Extract_By_Function_Out/results.tsv
83
- ```
84
-
85
- ---
86
-
87
- ## Output
88
-
89
- The tool generates a tab-delimited output file with the following columns:
90
-
91
- 1. **Sample:** Name of the processed `.tsv` file.
92
- 2. **Taxa:** Genus-level taxonomic assignment extracted from the `Lineage` column.
93
- 3. **Proportion:** Proportion of matches to the given functional ID within the sample.
94
-
95
- Example output:
96
-
97
- ```
98
- Function ID: GO:0002
99
- Sample Taxa Reads Assigned (Function) Proportion (Function) Proportion (Total Reads)
100
- PN0536_0003_S83_Final_Contig.tsv Gordonibacter 60788 0.075 0.002
101
- PN0536_0003_S83_Final_Contig.tsv Streptomyces 115671 0.142 0.004
102
- PN0536_0003_S83_Final_Contig.tsv unknown 80890 0.099 0.003
103
- PN0536_0003_S83_Final_Contig.tsv Clostridium 51018 0.063 0.002
104
- PN0536_0003_S83_Final_Contig.tsv Lactobacillus 149909 0.184 0.005
105
- PN0536_0003_S83_Final_Contig.tsv Limosilactobacillus 79694 0.098 0.003
106
- ```
107
-
108
- ---
109
-
110
- ## Implementation Details
111
-
112
- ### Workflow
113
-
114
- 1. The script reads `.tsv` files from the specified directory.
115
- 2. For each file, it searches for occurrences of the given functional ID within specific columns.
116
- 3. Matches are associated with genus-level taxonomic information extracted from the `Lineage` column.
117
- 4. Taxa proportions are calculated and saved to the output file.
118
-
119
- ### Large File Handling (Might be a failure point)
120
-
121
- The script uses `csv.field_size_limit` to handle exceptionally large `.tsv` files.
122
-
123
- ---
124
-
125
- ## Future Plans
126
-
127
- - Add support for additional file formats (e.g., `.csv`, `.txt`).
128
- - Expand functionality for more complex taxonomic and functional analyses.
129
- ---
130
-
File without changes
File without changes