virp 1.0.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
virp-1.0.0/LICENSE ADDED
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2024 Kedar Hippalgaonkar's Materials by Design Lab
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
virp-1.0.0/PKG-INFO ADDED
@@ -0,0 +1,85 @@
1
+ Metadata-Version: 2.2
2
+ Name: virp
3
+ Version: 1.0.0
4
+ Summary: VIRtual cell generation by Permutation
5
+ Author-email: Andy Paul Chen <la.vache.qui.vit@gmail.com>
6
+ License: MIT License
7
+
8
+ Copyright (c) 2024 Kedar Hippalgaonkar's Materials by Design Lab
9
+
10
+ Permission is hereby granted, free of charge, to any person obtaining a copy
11
+ of this software and associated documentation files (the "Software"), to deal
12
+ in the Software without restriction, including without limitation the rights
13
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
14
+ copies of the Software, and to permit persons to whom the Software is
15
+ furnished to do so, subject to the following conditions:
16
+
17
+ The above copyright notice and this permission notice shall be included in all
18
+ copies or substantial portions of the Software.
19
+
20
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
21
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
22
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
23
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
24
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
25
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
26
+ SOFTWARE.
27
+
28
+ Keywords: disordered,virtual cell,cif,partial occupancy
29
+ Classifier: Programming Language :: Python :: 3.9
30
+ Classifier: License :: OSI Approved :: MIT License
31
+ Classifier: Operating System :: OS Independent
32
+ Description-Content-Type: text/markdown
33
+ License-File: LICENSE
34
+ Requires-Dist: pymatgen
35
+ Requires-Dist: chgnet
36
+ Requires-Dist: matgl==1.0.0
37
+ Requires-Dist: dgl==1.1.2
38
+ Requires-Dist: poshcar
39
+
40
+ <img src="graphics/virpbanner.png" width="870">
41
+
42
+ # `virp`: VIRtual cell generation by Permutation
43
+ `virp` is a code for the fast generation of a virtual cell from a crystal structure (in CIF format) containing site disorder. It is named after Singapore's first superhero, VR Man, whose superpower is "Virping". The show was a flop, but we are still proud of him.
44
+
45
+ This project is inspired by the `Supercell` code of Okhotnikov, Charpentier and Cadars (<i>J. Cheminform. <b>8</b>, 17</i>), which formed the basis of our fast virtual cell generation algorithm, as well as the `aflow++` framework (<i>Comput. Mater. Sci. <b>217</b>, 111889</i>), for the statistical postprocessing of materials properties.
46
+
47
+ ## Theory
48
+ (To be updated!)
49
+
50
+ ## Requirements
51
+ `pymatgen`, `chgnet`, and `matgl` (`matgl==1.0.0`; `dgl==1.1.2`)<br>
52
+ __Optional__: You can also use git for the fancy installation. Otherwise, downloading the .py file will do.
53
+
54
+ ## Installation
55
+ `pip install git+https://github.com/andypaulchen/virp.git`<br>
56
+ Update to latest release: uninstall and re-install
57
+
58
+ ## Building a database
59
+ The root directory has a folder (`session`) which holds the python scripts which build a library of virtual cells (`generate.py`) and postprocessing scripts (`connectivity.py` and `properties.py`). After each script is run, the results are saved as `.csv` files.
60
+
61
+ 1. To prepare for a session, copy the `session` folder in your workspace and place the `.cif` files you want to process (make virtual cells + postprocessing) in the subfolder `_disordered_cifs`. Feel free to rename `session` folder to something more identifiable
62
+
63
+ 2. Run `generate.py` to create a supercell and (by default) 400 virtual cells.
64
+ - after this step, a structure subfolder (e.g. `structure`) is created in `session` for each `structure.cif` file in `_disordered_cifs`, with the same name. Inside this folder is a supercell CIF and folders for structure-optimized (`stropt`) and non-structure-optimized virtual cells (`no_stropt`). The details of this run is recorded in `virp_session_summary.csv`.
65
+
66
+ 3. Run `connectivity.py` for atomic connectivity post-processing
67
+ - after this step, the results are written to `connectivity.csv` and `scatterplot.png` under `stropt` and `no_stropt`.
68
+
69
+ 4. Run `properties.py` to predict materials properties. This is performed on `stropt` subfolders only.
70
+ - after this step, the results are written to `virtual_properties.csv` in the `structure` subfolder.
71
+
72
+ In summary, this is what a session looks like after all three routines have completed:
73
+
74
+ <img src="graphics/operation.png" width="870">
75
+
76
+ ## Versions and changelog
77
+ `v0.1.1`: first workable code, with function to generate a virtual cell. <br>
78
+ `v0.2.1`: added enumeration function <br>
79
+ `v0.2.2`: enumeration can be imported now (fix) <br>
80
+ `v0.3.0`: you can now make a batch of virtual cells<br>
81
+ `v0.4.3`: added tools to build a database
82
+
83
+ ## Debugging and support
84
+ The `virp` code has been tested on a limited number of platforms, so far Windows and Linux. If you are running into any problems during operation, please hound me (Andy Paul Chen) at la.vache.qui.vit(at)gmail.com, and I will try my best to help.
85
+
virp-1.0.0/README.md ADDED
@@ -0,0 +1,46 @@
1
+ <img src="graphics/virpbanner.png" width="870">
2
+
3
+ # `virp`: VIRtual cell generation by Permutation
4
+ `virp` is a code for the fast generation of a virtual cell from a crystal structure (in CIF format) containing site disorder. It is named after Singapore's first superhero, VR Man, whose superpower is "Virping". The show was a flop, but we are still proud of him.
5
+
6
+ This project is inspired by the `Supercell` code of Okhotnikov, Charpentier and Cadars (<i>J. Cheminform. <b>8</b>, 17</i>), which formed the basis of our fast virtual cell generation algorithm, as well as the `aflow++` framework (<i>Comput. Mater. Sci. <b>217</b>, 111889</i>), for the statistical postprocessing of materials properties.
7
+
8
+ ## Theory
9
+ (To be updated!)
10
+
11
+ ## Requirements
12
+ `pymatgen`, `chgnet`, and `matgl` (`matgl==1.0.0`; `dgl==1.1.2`)<br>
13
+ __Optional__: You can also use git for the fancy installation. Otherwise, downloading the .py file will do.
14
+
15
+ ## Installation
16
+ `pip install git+https://github.com/andypaulchen/virp.git`<br>
17
+ Update to latest release: uninstall and re-install
18
+
19
+ ## Building a database
20
+ The root directory has a folder (`session`) which holds the python scripts which build a library of virtual cells (`generate.py`) and postprocessing scripts (`connectivity.py` and `properties.py`). After each script is run, the results are saved as `.csv` files.
21
+
22
+ 1. To prepare for a session, copy the `session` folder in your workspace and place the `.cif` files you want to process (make virtual cells + postprocessing) in the subfolder `_disordered_cifs`. Feel free to rename `session` folder to something more identifiable
23
+
24
+ 2. Run `generate.py` to create a supercell and (by default) 400 virtual cells.
25
+ - after this step, a structure subfolder (e.g. `structure`) is created in `session` for each `structure.cif` file in `_disordered_cifs`, with the same name. Inside this folder is a supercell CIF and folders for structure-optimized (`stropt`) and non-structure-optimized virtual cells (`no_stropt`). The details of this run is recorded in `virp_session_summary.csv`.
26
+
27
+ 3. Run `connectivity.py` for atomic connectivity post-processing
28
+ - after this step, the results are written to `connectivity.csv` and `scatterplot.png` under `stropt` and `no_stropt`.
29
+
30
+ 4. Run `properties.py` to predict materials properties. This is performed on `stropt` subfolders only.
31
+ - after this step, the results are written to `virtual_properties.csv` in the `structure` subfolder.
32
+
33
+ In summary, this is what a session looks like after all three routines have completed:
34
+
35
+ <img src="graphics/operation.png" width="870">
36
+
37
+ ## Versions and changelog
38
+ `v0.1.1`: first workable code, with function to generate a virtual cell. <br>
39
+ `v0.2.1`: added enumeration function <br>
40
+ `v0.2.2`: enumeration can be imported now (fix) <br>
41
+ `v0.3.0`: you can now make a batch of virtual cells<br>
42
+ `v0.4.3`: added tools to build a database
43
+
44
+ ## Debugging and support
45
+ The `virp` code has been tested on a limited number of platforms, so far Windows and Linux. If you are running into any problems during operation, please hound me (Andy Paul Chen) at la.vache.qui.vit(at)gmail.com, and I will try my best to help.
46
+
@@ -0,0 +1,31 @@
1
+ [build-system]
2
+ requires = ["setuptools>=61.0", "wheel", "pip<24.1"]
3
+ build-backend = "setuptools.build_meta"
4
+
5
+ [project]
6
+ name = "virp"
7
+ version = "1.0.0"
8
+ description = "VIRtual cell generation by Permutation"
9
+ readme = "README.md"
10
+ license = {file = "LICENSE"}
11
+ authors = [
12
+ {name = "Andy Paul Chen", email = "la.vache.qui.vit@gmail.com"}
13
+ ]
14
+ dependencies = [
15
+ "pymatgen",
16
+ "chgnet",
17
+ "matgl==1.0.0",
18
+ "dgl==1.1.2",
19
+ "poshcar"
20
+ ]
21
+ keywords = [
22
+ "disordered",
23
+ "virtual cell",
24
+ "cif",
25
+ "partial occupancy"
26
+ ]
27
+ classifiers = [
28
+ "Programming Language :: Python :: 3.9",
29
+ "License :: OSI Approved :: MIT License",
30
+ "Operating System :: OS Independent"
31
+ ]
virp-1.0.0/setup.cfg ADDED
@@ -0,0 +1,4 @@
1
+ [egg_info]
2
+ tag_build =
3
+ tag_date = 0
4
+
virp-1.0.0/setup.py ADDED
@@ -0,0 +1,14 @@
1
+ from setuptools import setup, find_packages
2
+
3
+ setup(
4
+ name="virp",
5
+ version="1.0.0",
6
+ packages=find_packages(),
7
+ install_requires=[
8
+ "pymatgen",
9
+ "chgnet",
10
+ "matgl==1.0.0",
11
+ "dgl==1.1.2",
12
+ "poshcar"
13
+ ],
14
+ )
@@ -0,0 +1,16 @@
1
+ # Copyright (c) 2024, Kedar Hippalgaonkar's Materials by Design Lab
2
+ # Distributed under the terms of the MIT License.
3
+
4
+ """
5
+ Virp (VIRtual cell generation by Permutation) is a code for generating a virtual cell from a site-disordered .cif crystal file
6
+ """
7
+
8
+ from .main import *
9
+ from .enumerate import *
10
+
11
+ __copyright__ = "Kedar Hippalgaonkar's Materials by Design Lab"
12
+ __version__ = "1.0.0"
13
+ __maintainer__ = "Andy Paul Chen"
14
+ __email__ = "la.vache.qui.vit@gmail.com"
15
+ __status__ = "Development"
16
+ __date__ = "10 October 2024"
@@ -0,0 +1,113 @@
1
+ # database.py
2
+
3
+ # External Imports
4
+ from pymatgen.io.cif import CifParser # write pymatgen structure to cif
5
+ from pathlib import Path
6
+ from tqdm import tqdm
7
+ import os
8
+
9
+ def DisorderQuery(folder_path):
10
+ """
11
+ Process all CIF files in a folder to check for partial occupancy.
12
+ Displays a progress bar and summary statistics.
13
+
14
+ Parameters:
15
+ -----------
16
+ folder_path : str
17
+ Path to the folder containing CIF files
18
+ threshold : float, optional
19
+ Occupancy threshold for checking partial occupancy
20
+
21
+ Returns:
22
+ --------
23
+ dict
24
+ Dictionary with CIF filenames as keys and their analysis results as values
25
+ """
26
+ folder = Path(folder_path)
27
+
28
+ if not folder.is_dir():
29
+ raise NotADirectoryError(f"Folder not found: {folder_path}")
30
+
31
+ results = {}
32
+ cif_files = list(folder.glob("*.cif"))
33
+
34
+ # Initialize counters
35
+ total_files = len(cif_files)
36
+ files_with_partial = 0
37
+ files_without_partial = 0
38
+ error_files = 0
39
+
40
+ # Process each CIF file with progress bar
41
+ for cif_file in tqdm(cif_files, desc="Processing CIF files", unit="file"):
42
+ try:
43
+ result = is_site_disordered(str(cif_file))
44
+ results[cif_file.name] = result
45
+
46
+ # Update counters silently
47
+ if result["has_partial"]:
48
+ files_with_partial += 1
49
+ else:
50
+ files_without_partial += 1
51
+
52
+ except Exception as e:
53
+ error_files += 1
54
+ results[cif_file.name] = {"error": str(e)}
55
+
56
+ # Print final summary statistics
57
+ print("\nSummary Statistics:")
58
+ print("-" * 50)
59
+ print(f"Total CIF files processed: {total_files}")
60
+ print(f"Files with partial occupancy: {files_with_partial} ({files_with_partial/total_files*100:.1f}%)")
61
+ print(f"Files without partial occupancy: {files_without_partial} ({files_without_partial/total_files*100:.1f}%)")
62
+ if error_files > 0:
63
+ print(f"Files with errors: {error_files} ({error_files/total_files*100:.1f}%)")
64
+
65
+ return results
66
+
67
+
68
+
69
+ def is_SiteDisordered(cif_path):
70
+ """
71
+ Check if a CIF file contains sites with partial occupancy.
72
+
73
+ Parameters:
74
+ -----------
75
+ cif_path : str
76
+ Path to the CIF file
77
+ threshold : float, optional
78
+ Occupancy threshold below which a site is considered partially occupied
79
+ Default is 1.0 (fully occupied)
80
+
81
+ Returns:
82
+ --------
83
+ dict
84
+ Dictionary containing:
85
+ - has_partial: bool, whether partial occupancy was found
86
+ - partial_sites: list of tuples (site index, species, occupancy)
87
+ """
88
+ # Verify file exists
89
+ if not os.path.exists(cif_path):
90
+ raise FileNotFoundError(f"CIF file not found: {cif_path}")
91
+
92
+ # Parse the CIF file
93
+ parser = CifParser(cif_path)
94
+ structure = parser.get_structures()[0]
95
+
96
+ # Initialize results
97
+ partial_sites = []
98
+
99
+ # Check each site in the structure
100
+ for i, site in enumerate(structure.sites):
101
+ species_dict = site.species.as_dict()
102
+
103
+ # Check occupancy for each species on the site
104
+ for element, occupancy in species_dict.items():
105
+ if occupancy < 1.0:
106
+ partial_sites.append((i, element, occupancy))
107
+
108
+ result = {
109
+ "has_partial": len(partial_sites) > 0,
110
+ "partial_sites": partial_sites
111
+ }
112
+
113
+ return result
@@ -0,0 +1,160 @@
1
+ # enumerate.py: counts possible permutations and combinations for atom filling in disordered sites
2
+
3
+ from itertools import product
4
+ from math import factorial, prod
5
+ import numpy as np
6
+ import re
7
+
8
+ def format_integer(num, prec = 6):
9
+ return np.format_float_scientific(num, precision=prec) if num >= 10**prec else str(num)
10
+
11
+
12
+ def discretize_floats(arr):
13
+ # Store possible discretizations for each float
14
+ discretizations = []
15
+
16
+ for num in arr:
17
+ if num % 1 == 0.5: # Equidistant case
18
+ lower = int(num // 1) # Round down
19
+ upper = lower + 1 # Round up
20
+ discretizations.append([lower, upper])
21
+ else:
22
+ discretizations.append([round(num)]) # Standard rounding
23
+
24
+ # Generate all combinations of discretizations
25
+ all_discretizations = [list(discretization) for discretization in product(*discretizations)]
26
+
27
+ return all_discretizations
28
+
29
+
30
+ def remove_duplicate_sublists(lst):
31
+ seen = set()
32
+ unique_sublists = []
33
+ for sublist in lst:
34
+ sublist_tuple = tuple(sublist)
35
+ if sublist_tuple not in seen:
36
+ seen.add(sublist_tuple)
37
+ unique_sublists.append(sublist)
38
+ return unique_sublists
39
+
40
+
41
+ def enumerate_site(N, compositions, verbose = True):
42
+ # Enumerate combination of enumerations by disordered site
43
+ # N: number of sites
44
+ # compositions: [float], partition fractions adding up to < 1
45
+ if sum(compositions) > 1: print("Error: Compositions add up to more than 100%: ", compositions) # This no make sense (failsafe)
46
+ else:
47
+ if sum(compositions) < 1: compositions.append(1-sum(compositions)) # include vacancies in permutation
48
+ partitions = []
49
+ for i in range(len(compositions)):
50
+ partitions.append(sum(compositions[:i+1]))
51
+ partN = [i*N for i in partitions]
52
+
53
+ # initialize snapping
54
+ total_combinations = factorial(N)
55
+ if verbose: print("- Raw permutations: ", format_integer(total_combinations), "(", N, "!)")
56
+
57
+ # discretize floats
58
+ all_snaps = discretize_floats(partN)
59
+ # assign at least 1 atom per element
60
+ for snap in all_snaps:
61
+ for index in range(len(snap)):
62
+ if index > 0:
63
+ if snap[index] == snap[index-1]:
64
+ if verbose and (snap[index] + 1 >= N): print("Error: Choose a bigger supercell!")
65
+ else: snap[index] += 1
66
+ # remove duplicates in all_snaps
67
+ all_snaps = remove_duplicate_sublists(all_snaps)
68
+
69
+ # for each snap, calculate number of combinations
70
+ allcombinations = 0
71
+ for snap in all_snaps:
72
+ if verbose: print("- Snap: ", snap)
73
+ combination = total_combinations
74
+ for index in range(len(snap)):
75
+ if index == 0: n = snap[index]
76
+ else: n = snap[index]-snap[index-1]
77
+ combination /= factorial(n)
78
+ thiscombination = int(combination)
79
+ allcombinations += thiscombination
80
+ if verbose: print("- No. of combinations: ", format_integer(thiscombination))
81
+
82
+ return all_snaps, allcombinations
83
+
84
+
85
+ def get_site_combination(edit_block, edit_name):
86
+ # Auxiliary function which, outside of the enumerate structure routine, will make no sense whatsoever
87
+
88
+ # 1. What are the unique elements and occupancies?
89
+ atomoccpairslist = []
90
+ for evalline in edit_block:
91
+ # Split each line into components (using split will automatically handle whitespaces)
92
+ parts = evalline.split()
93
+ atomoccpair = (parts[0], float(parts[-1]))
94
+ if atomoccpair not in atomoccpairslist:
95
+ atomoccpairslist.append(atomoccpair)
96
+
97
+ # Display specifications
98
+ print("Disordered site name: ", edit_name)
99
+ numberoflines = len(edit_block)
100
+ print("- Number of sites in supercell: ", numberoflines)
101
+ print("- Element and occupancy: ", atomoccpairslist) # The number of elements in this site = N
102
+ proportions = [t[1] for t in atomoccpairslist]
103
+ combinations = enumerate_site(numberoflines, proportions)[1]
104
+
105
+ return combinations
106
+
107
+
108
+ def enumerate_structure(input_file):
109
+ # Given a SUPERCELL .cif structure, return total possible virtual cells,
110
+ # disregarding symmetry equivalence
111
+ print("Input supercell .cif file: ", input_file)
112
+
113
+ # Updated regex pattern to capture the second string and the last number
114
+ pattern = re.compile(r'\s*\S+\s+(\S+)\s+1\s+[0-9]+\.[0-9]+\s+[0-9]+\.[0-9]+\s+[0-9]+\.[0-9]+\s+([0-9]+\.[0-9]+)')
115
+ product_list = [] # list of permutations to include in
116
+
117
+ # Open the input file to read and the output file to write
118
+ with open(input_file, 'r') as infile:
119
+ # Declare edit space (as in permutative fill, but without updating)
120
+ edit_active = False # is thisline in an editing block?
121
+ edit_block = [] # array to store lines in an editing block
122
+ edit_name = "" # stores the site which forms the edit block
123
+
124
+ for thisline in infile: # scan through the file
125
+ # Check if the line matches the pattern
126
+ match = pattern.match(thisline)
127
+
128
+ if match: # we have reached the coordinate block of the .cif file
129
+ # Extract the site name and last number from the match
130
+ second_string = match.group(1) # This will give you 'Ca1'
131
+ last_number = float(match.group(2)) # The last number
132
+
133
+ # Decision block
134
+ if last_number < 1.0: # Check if the last number is less than 1.0: partial occupancy site
135
+ if not edit_active: # if first line in an edit block
136
+ edit_active = True # switch on editing mode
137
+ edit_name = second_string # What site is being edited
138
+ else:
139
+ if not edit_name == second_string: # if a different site is being considered
140
+ product_list.append(get_site_combination(edit_block, edit_name)) # get combinations for site
141
+ # Re-initialize edit parameters
142
+ edit_block = [] # array to store lines in an editing block
143
+ edit_active = True
144
+ edit_name = second_string
145
+
146
+ else: # if no longer partial occupancy site
147
+ if edit_active:
148
+ product_list.append(get_site_combination(edit_block, edit_name)) # get combinations for site
149
+ # Re-initialize edit parameters
150
+ edit_active = False # switch off edit mode
151
+ edit_block = [] # array to store lines in an editing block
152
+ edit_name = "" # stores the site which forms the edit block
153
+
154
+ if edit_active:
155
+ # Write the line to the edit block
156
+ edit_block.append(thisline)
157
+
158
+ totalcombinations = prod(product_list)
159
+ print("Total number of combinations for", input_file, ": ", format_integer(totalcombinations))
160
+ return totalcombinations
@@ -0,0 +1,320 @@
1
+ # main.py
2
+
3
+ # External Imports
4
+ from pymatgen.core.structure import Structure
5
+ from chgnet.model import StructOptimizer
6
+ from itertools import product
7
+ import numpy as np
8
+ import pandas as pd
9
+ import warnings
10
+ import random
11
+ import math
12
+ import re
13
+ import os
14
+
15
+
16
+ def CIFSupercell (inputcif, outputcif, supercellsize):
17
+ # inputcif, outputcif: path to cif file
18
+ # supercellsize: vector of 3 integers
19
+
20
+ # Load the structure from a CIF file
21
+ structure = Structure.from_file(inputcif)
22
+
23
+ # Define the scaling matrix for the supercell
24
+ # For example, [2, 0, 0], [0, 2, 0], [0, 0, 2] creates a 2x2x2 supercell
25
+ scaling_matrix = [[supercellsize[0], 0, 0],
26
+ [0, supercellsize[1], 0],
27
+ [0, 0, supercellsize[2]]]
28
+
29
+ # Create the supercell
30
+ structure.make_supercell(scaling_matrix)
31
+
32
+ # Save the supercell to a new CIF file (optional)
33
+ structure.to(fmt="cif", filename=outputcif)
34
+ print("Supercell created and saved as ", outputcif)
35
+
36
+
37
+ def round_with_tie_breaker(n):
38
+ # Separate the fractional and integer parts
39
+ fractional_part, integer_part = math.modf(n)
40
+
41
+ # Check if the fractional part is 0.5
42
+ if abs(fractional_part) == 0.5:
43
+ # Randomly choose to round down or up
44
+ return int(integer_part) + random.choice([0, 1])
45
+ else:
46
+ # Regular rounding for non 0.5 cases
47
+ return round(n)
48
+
49
+
50
+ def ShuffleOccupiedSites (outfile, edit_block, edit_name):
51
+ # Auxiliary function which, outside of the permutative fill routine, will make no sense whatsoever
52
+
53
+ # 1. What are the unique elements and occupancies?
54
+ atomoccpairslist = []
55
+ for evalline in edit_block:
56
+ # Split each line into components (using split will automatically handle whitespaces)
57
+ parts = evalline.split()
58
+ atomoccpair = (parts[0], float(parts[-1]))
59
+ if atomoccpair not in atomoccpairslist:
60
+ atomoccpairslist.append(atomoccpair)
61
+
62
+ # Display specifications
63
+ print("Disordered site name: ", edit_name)
64
+ numberofelements = len(atomoccpairslist)
65
+ print("- Number of elements in this site: ", numberofelements) # The number of elements in this site = N
66
+
67
+ # Keep every Nth line in the edit block
68
+ if numberofelements > 1: edit_block = edit_block[::numberofelements]
69
+
70
+ # Randomly shuffle the list
71
+ random.shuffle(edit_block)
72
+
73
+ # Assign atoms based on proportion in atomoccpairslist
74
+ numberoflines = len(edit_block)
75
+ print("- Number of sites in supercell: ", numberoflines)
76
+
77
+ atomassignmentlist_float = []
78
+ assignment_cumulative = 0
79
+ assignment_cumulative_int = 0
80
+ for atomoccpair in atomoccpairslist:
81
+ # evaluate how many atoms to assign to element in question
82
+ atomassignment_float = atomoccpair[1]*numberoflines
83
+ assignment_cumulative += atomassignment_float
84
+ assignment_int = max(round_with_tie_breaker(assignment_cumulative)-assignment_cumulative_int,1) # assign at least 1 atom
85
+ assignment_cumulative_int += assignment_int
86
+ # tuples for display
87
+ atomassignmentlist_float.append((atomoccpair[0], atomassignment_float, assignment_int))
88
+
89
+ print("- Atoms and site assignment (float/rounded): ", atomassignmentlist_float)
90
+ print("- No of filled sites: ", assignment_cumulative_int,"/",len(edit_block))
91
+ edit_block = edit_block[:assignment_cumulative_int]
92
+
93
+ # Implement the atom-site assignment in-text
94
+ pointer = 0 # line-by line pointer for edit_block rows
95
+ for this_element in atomassignmentlist_float:
96
+ element_name = this_element[0]
97
+ no_atoms = this_element[2]
98
+ for i in range(no_atoms):
99
+ edit_block[pointer] = re.sub(r'^(\s*)([^\s]+)', r'\1' + element_name, edit_block[pointer])
100
+ pointer += 1
101
+
102
+ # Change every occupancy to 1.0
103
+ edit_block = [re.sub(r'([0-9]+\.[0-9]+)\s*$', '1.0', line) + '\n' for line in edit_block]
104
+
105
+ for writeline in edit_block: outfile.write(writeline)
106
+
107
+
108
+ def PermutativeFill(input_file, output_file):
109
+ # Updated regex pattern to capture the second string and the last number
110
+ pattern = re.compile(r'\s*\S+\s+(\S+)\s+1\s+[0-9]+\.[0-9]+\s+[0-9]+\.[0-9]+\s+[0-9]+\.[0-9]+\s+([0-9]+\.[0-9]+)')
111
+
112
+ # Open the input file to read and the output file to write
113
+ with open(input_file, 'r') as infile, open(output_file, 'w') as outfile:
114
+ # Declare edit space (a series of lines where permutative fill takes place)
115
+ edit_active = False # is thisline in an editing block?
116
+ edit_block = [] # array to store lines in an editing block
117
+ edit_name = "" # stores the site which forms the edit block
118
+
119
+ for thisline in infile: # scan through the file
120
+ # Check if the line matches the pattern
121
+ match = pattern.match(thisline)
122
+
123
+ if match: # we have reached the coordinate block of the .cif file
124
+ # Extract the site name and last number from the match
125
+ second_string = match.group(1) # This will give you 'Ca1'
126
+ last_number = float(match.group(2)) # The last number
127
+
128
+ # Decision block
129
+
130
+ if last_number < 1.0: # Check if the last number is less than 1.0: partial occupancy site
131
+ if not edit_active: # if first line in an edit block
132
+ edit_active = True # switch on editing mode
133
+ edit_name = second_string # What site is being edited
134
+ else:
135
+ if not edit_name == second_string: # if a different site is being considered
136
+ ShuffleOccupiedSites(outfile, edit_block, edit_name) # WRITE EDITING BLOCK TO FILE; this also resets it to []
137
+ # Re-initialize edit parameters
138
+ edit_block = [] # array to store lines in an editing block
139
+ edit_active = True
140
+ edit_name = second_string
141
+
142
+ else: # if no longer partial occupancy site
143
+ if edit_active:
144
+ ShuffleOccupiedSites(outfile, edit_block, edit_name) # WRITE EDITING BLOCK TO FILE
145
+ # Re-initialize edit parameters
146
+ edit_active = False # switch off edit mode
147
+ edit_block = [] # array to store lines in an editing block
148
+ edit_name = "" # stores the site which forms the edit block
149
+
150
+ # Execution block
151
+
152
+ if edit_active:
153
+ # Write the line to the edit block
154
+ edit_block.append(thisline)
155
+
156
+ else: # edit mode is not active
157
+ # Write the thisline to the output file
158
+ outfile.write(thisline)
159
+
160
+ else: # other lines we are not bothered with
161
+ outfile.write(thisline)
162
+ if edit_active: # we have reached the end of the coordinate block
163
+ edit_active = False # switch off edit mode
164
+
165
+ # WRITE SEQUENCE
166
+ ShuffleOccupiedSites(outfile, edit_block, edit_name)
167
+ # Re-initialize edit parameters
168
+ edit_block = [] # array to store lines in an editing block
169
+ edit_name = "" # stores the site which forms the edit block
170
+
171
+
172
+ def SampleVirtualCells(input_cif, supercell, sample_size=400):
173
+ """
174
+ Given a disordered .cif file, create an output folder
175
+ containing a number (sample_size) of virtual cells
176
+
177
+ Args:
178
+ input_cif (str): Path to .cif (disordered)
179
+ supercell [int,int,int]: multiplicity of supercell
180
+ sample_size (int): Number of virtual cells to generate (default is 400)
181
+
182
+ Returns:
183
+ void
184
+ """
185
+ # Init CHGNET optimizer
186
+ relaxer = StructOptimizer()
187
+
188
+ # Suppress warnings in this block
189
+ with warnings.catch_warnings():
190
+ warnings.simplefilter("ignore")
191
+
192
+ # Make output folder directory
193
+ fname = os.path.splitext(os.path.basename(input_cif))[0]
194
+ os.makedirs(fname, exist_ok=True) # `exist_ok=True` avoids errors if the directory exists.
195
+ print(f"Directory created at: {fname}")
196
+
197
+ header = os.path.join(fname,fname)
198
+ sc_file = header+"_supercell.cif"
199
+
200
+ # Make the supercell
201
+ CIFSupercell (input_cif, sc_file, supercell)
202
+
203
+ # Create target folders if they don't exist
204
+ stropt_path = os.path.join(fname,"stropt")
205
+ no_stropt_path = os.path.join(fname,"no_stropt")
206
+ os.makedirs(stropt_path, exist_ok=True) # structure-optimized cells
207
+ os.makedirs(no_stropt_path, exist_ok=True) # non-structure-optimized cells
208
+
209
+ # Execution
210
+ for i in range(sample_size):
211
+ # Permutative fill only, no structure optimization
212
+ print("Generating virtual cell #", i, ":")
213
+ pfill_file_name = fname+"_virtual_"+str(i)+".cif"
214
+ pfill_file = os.path.join(no_stropt_path,pfill_file_name)
215
+ PermutativeFill(sc_file, pfill_file)
216
+
217
+ # Relax
218
+ structure = Structure.from_file(pfill_file)
219
+ result = relaxer.relax(structure, verbose=False)
220
+ stropt_file_name = fname+"_virtual_"+str(i)+"_stropt.cif"
221
+ stropt_file = os.path.join(stropt_path,stropt_file_name)
222
+ result['final_structure'].to(stropt_file)
223
+
224
+ with open(os.path.join(fname,"_JOBDONE"), 'w') as file: pass # make an empty file signalling completion
225
+ print("All cells generated (see _JOBDONE file).")
226
+
227
+
228
+ def SupercellSize(input_cif, minsize=15.0):
229
+ """
230
+ Given a disordered .cif file, decide how big the
231
+ supercell should be (works best for orthogonal cifs)
232
+
233
+ Args:
234
+ input_cif (str): Path to .cif (disordered)
235
+ minsize (float): minimum tolerated distance between
236
+ lattice points in one direction
237
+
238
+ Returns:
239
+ array of 3 integers denoting supercell multiplicity
240
+ """
241
+ # init sc_size array, warning
242
+ sc_size = [0,0,0]
243
+ warning = False
244
+
245
+ # Load the .cif file
246
+ structure = Structure.from_file(input_cif)
247
+
248
+ # Get the lattice vectors
249
+ lattice = structure.lattice
250
+ new_lattice = []
251
+
252
+ # Execution
253
+ for i in range(3):
254
+ uc_length = np.linalg.norm(lattice.matrix[i])
255
+ sc_size[i] = math.ceil(minsize/uc_length)
256
+ new_lattice.append(lattice.matrix[i]*sc_size[i])
257
+
258
+ # Generate all lattice points for one unit cell
259
+ lattice_points = [np.dot([i, j, k], new_lattice) for i, j, k in product([0, 1], repeat=3)]
260
+ # Calculate all pairwise distances
261
+ distances = []
262
+ for i, p1 in enumerate(lattice_points):
263
+ for j, p2 in enumerate(lattice_points):
264
+ if i < j: # Avoid duplicate pairs
265
+ distances.append(np.linalg.norm(p1 - p2))
266
+
267
+ # Find the shortest distance
268
+ shortest_lattice_distance = min(distances)
269
+
270
+ # Check if shortest distance between lattice points is under minsize
271
+ print(f"The shortest distance between lattice points is: {shortest_lattice_distance:.5f} Å")
272
+ if shortest_lattice_distance < minsize:
273
+ print("Warning: lattice points still close together for supercell; check orthogonality!")
274
+ warning = True
275
+ print(f"Supercell multiplicity: {sc_size}")
276
+
277
+ return sc_size, warning
278
+
279
+
280
+ def Session(folder_path = "_disordered_cifs", mindist = 15, no_of_samples = 400):
281
+ # init DataFrame to store results
282
+ data = []
283
+
284
+ # Loop through all .cif files in the folder
285
+ for filename in os.listdir(folder_path):
286
+ if filename.endswith(".cif"): # Check if the file has a .cif extension
287
+ file_path = os.path.join(folder_path, filename)
288
+ print(f"Processing .cif file: {file_path}")
289
+
290
+ try:
291
+ # Calculate preferred supercell size
292
+ sc_size, warning = SupercellSize(file_path, minsize=mindist)
293
+
294
+ # Generate virtual cell samples
295
+ SampleVirtualCells(file_path, sc_size, sample_size=no_of_samples)
296
+
297
+ # Extract metadata: chemical formula
298
+ structure = Structure.from_file(file_path)
299
+ formula = structure.composition.reduced_formula
300
+
301
+ # Append results to the data list
302
+ data.append({
303
+ "filename": filename,
304
+ "folder": folder_path,
305
+ "formula": formula,
306
+ "supercell size": sc_size,
307
+ "sample size": no_of_samples,
308
+ "lattice spacing warning": warning
309
+ })
310
+
311
+ except Exception as e:
312
+ print(f"Error processing {file_path}: {e}")
313
+
314
+ # Create a DataFrame
315
+ df = pd.DataFrame(data)
316
+
317
+ # Save the DataFrame to a CSV file
318
+ output_file = "virp_session_summary.csv"
319
+ df.to_csv(output_file)
320
+ print(f"Results saved to {output_file}")
@@ -0,0 +1,120 @@
1
+ # matprop.py
2
+
3
+ from pymatgen.core.structure import Structure
4
+ import os
5
+ import csv
6
+ import torch
7
+ import matgl
8
+ from chgnet.model.model import CHGNet
9
+
10
+ def VirtualCellProperties(folder_path, output_csv):
11
+ # To add: customise the set of properties to evaluate
12
+ """
13
+ Given a folder filled with virtual cells,
14
+ predict material properties for each virtual cell,
15
+ and write results in a .csv form
16
+
17
+ Args:
18
+ folder_path (str): Path to folder
19
+ output_csv (str): Path to .csv output
20
+
21
+ Returns:
22
+ void
23
+ """
24
+ # Load the MEGNet band gap model
25
+ bandgap_model = matgl.load_model("MEGNet-MP-2019.4.1-BandGap-mfi")
26
+
27
+ # Load the CHGNet model for total energy prediction
28
+ chgnet = CHGNet.load()
29
+
30
+ # Initialize data storage
31
+ data = []
32
+
33
+ for filename in os.listdir(folder_path):
34
+ if filename.endswith("stropt.cif"): # evaluate structure-optimized cells only
35
+ filepath = os.path.join(folder_path, filename)
36
+ try:
37
+ # Load the structure
38
+ structure = Structure.from_file(filepath)
39
+
40
+ # Predict total energy
41
+ total_energy = chgnet.predict_structure(structure)['e']
42
+
43
+ # Calculate density and convert to float
44
+ density = float(structure.density)
45
+
46
+ # Predict band gaps for different methods
47
+ bandgaps = {}
48
+ for i, method in ((0, "PBE"), (1, "GLLB-SC"), (2, "HSE"), (3, "SCAN")):
49
+ graph_attrs = torch.tensor([i])
50
+ bandgap = bandgap_model.predict_structure(structure=structure, state_attr=graph_attrs)
51
+ bandgaps[method] = float(bandgap)
52
+
53
+ # Append results to data
54
+ data.append({
55
+ "File": filename,
56
+ "Total Energy (eV)": total_energy,
57
+ "Density": density,
58
+ "PBE Bandgap (eV)": bandgaps["PBE"],
59
+ "GLLB-SC Bandgap (eV)": bandgaps["GLLB-SC"],
60
+ "HSE Bandgap (eV)": bandgaps["HSE"],
61
+ "SCAN Bandgap (eV)": bandgaps["SCAN"],
62
+ })
63
+
64
+ print(f"Processed: {filename}")
65
+
66
+ except Exception as e:
67
+ print(f"Error processing {filename}: {e}")
68
+
69
+ # Write results to CSV
70
+ with open(output_csv, mode='w', newline='') as csvfile:
71
+ fieldnames = ["File", "Total Energy (eV)", "Density", "PBE Bandgap (eV)", "GLLB-SC Bandgap (eV)", "HSE Bandgap (eV)", "SCAN Bandgap (eV)"]
72
+ writer = csv.DictWriter(csvfile, fieldnames=fieldnames)
73
+
74
+ writer.writeheader()
75
+ writer.writerows(data)
76
+
77
+ print(f"Results saved to {output_csv}")
78
+
79
+
80
+ import pandas as pd
81
+ import numpy as np
82
+
83
+ def ExpectationValues(csv_path, temperature):
84
+ """
85
+ Calculate Boltzmann-weighted expectation values for all numeric properties
86
+
87
+ Args:
88
+ csv_path (str): Path to CSV file
89
+ temperature (float): Temperature in Kelvin
90
+
91
+ Returns:
92
+ tuple: (DataFrame, dictionary of expectation values)
93
+ """
94
+ # Read the CSV file
95
+ df = pd.read_csv(csv_path)
96
+
97
+ # Boltzmann constant in eV/K = 0.00008617
98
+ k_B = 0.0000861733326
99
+
100
+ # Calculate weights using the Boltzmann distribution formula
101
+ df['weights'] = np.exp(-df['Total Energy (eV)']/(k_B * temperature))
102
+
103
+ # Calculate total weights
104
+ total_weights = df['weights'].sum()
105
+
106
+ # Dictionary to store expectation values
107
+ expectation_values = {}
108
+
109
+ # Get all numeric columns except 'Total Energy (eV)' and 'weights'
110
+ excluded_cols = ['File', 'Total Energy (eV)', 'weights']
111
+ numeric_cols = df.select_dtypes(include=[np.number]).columns
112
+ properties = [col for col in numeric_cols if col not in excluded_cols]
113
+
114
+ # Calculate weighted properties and their expectation values
115
+ for prop in properties:
116
+ weighted_col_name = f'weighted_{prop}'
117
+ df[weighted_col_name] = (df[prop] * df['weights']) / total_weights
118
+ expectation_values[prop] = df[weighted_col_name].sum()
119
+
120
+ return df, expectation_values
@@ -0,0 +1,85 @@
1
+ Metadata-Version: 2.2
2
+ Name: virp
3
+ Version: 1.0.0
4
+ Summary: VIRtual cell generation by Permutation
5
+ Author-email: Andy Paul Chen <la.vache.qui.vit@gmail.com>
6
+ License: MIT License
7
+
8
+ Copyright (c) 2024 Kedar Hippalgaonkar's Materials by Design Lab
9
+
10
+ Permission is hereby granted, free of charge, to any person obtaining a copy
11
+ of this software and associated documentation files (the "Software"), to deal
12
+ in the Software without restriction, including without limitation the rights
13
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
14
+ copies of the Software, and to permit persons to whom the Software is
15
+ furnished to do so, subject to the following conditions:
16
+
17
+ The above copyright notice and this permission notice shall be included in all
18
+ copies or substantial portions of the Software.
19
+
20
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
21
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
22
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
23
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
24
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
25
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
26
+ SOFTWARE.
27
+
28
+ Keywords: disordered,virtual cell,cif,partial occupancy
29
+ Classifier: Programming Language :: Python :: 3.9
30
+ Classifier: License :: OSI Approved :: MIT License
31
+ Classifier: Operating System :: OS Independent
32
+ Description-Content-Type: text/markdown
33
+ License-File: LICENSE
34
+ Requires-Dist: pymatgen
35
+ Requires-Dist: chgnet
36
+ Requires-Dist: matgl==1.0.0
37
+ Requires-Dist: dgl==1.1.2
38
+ Requires-Dist: poshcar
39
+
40
+ <img src="graphics/virpbanner.png" width="870">
41
+
42
+ # `virp`: VIRtual cell generation by Permutation
43
+ `virp` is a code for the fast generation of a virtual cell from a crystal structure (in CIF format) containing site disorder. It is named after Singapore's first superhero, VR Man, whose superpower is "Virping". The show was a flop, but we are still proud of him.
44
+
45
+ This project is inspired by the `Supercell` code of Okhotnikov, Charpentier and Cadars (<i>J. Cheminform. <b>8</b>, 17</i>), which formed the basis of our fast virtual cell generation algorithm, as well as the `aflow++` framework (<i>Comput. Mater. Sci. <b>217</b>, 111889</i>), for the statistical postprocessing of materials properties.
46
+
47
+ ## Theory
48
+ (To be updated!)
49
+
50
+ ## Requirements
51
+ `pymatgen`, `chgnet`, and `matgl` (`matgl==1.0.0`; `dgl==1.1.2`)<br>
52
+ __Optional__: You can also use git for the fancy installation. Otherwise, downloading the .py file will do.
53
+
54
+ ## Installation
55
+ `pip install git+https://github.com/andypaulchen/virp.git`<br>
56
+ Update to latest release: uninstall and re-install
57
+
58
+ ## Building a database
59
+ The root directory has a folder (`session`) which holds the python scripts which build a library of virtual cells (`generate.py`) and postprocessing scripts (`connectivity.py` and `properties.py`). After each script is run, the results are saved as `.csv` files.
60
+
61
+ 1. To prepare for a session, copy the `session` folder in your workspace and place the `.cif` files you want to process (make virtual cells + postprocessing) in the subfolder `_disordered_cifs`. Feel free to rename `session` folder to something more identifiable
62
+
63
+ 2. Run `generate.py` to create a supercell and (by default) 400 virtual cells.
64
+ - after this step, a structure subfolder (e.g. `structure`) is created in `session` for each `structure.cif` file in `_disordered_cifs`, with the same name. Inside this folder is a supercell CIF and folders for structure-optimized (`stropt`) and non-structure-optimized virtual cells (`no_stropt`). The details of this run is recorded in `virp_session_summary.csv`.
65
+
66
+ 3. Run `connectivity.py` for atomic connectivity post-processing
67
+ - after this step, the results are written to `connectivity.csv` and `scatterplot.png` under `stropt` and `no_stropt`.
68
+
69
+ 4. Run `properties.py` to predict materials properties. This is performed on `stropt` subfolders only.
70
+ - after this step, the results are written to `virtual_properties.csv` in the `structure` subfolder.
71
+
72
+ In summary, this is what a session looks like after all three routines have completed:
73
+
74
+ <img src="graphics/operation.png" width="870">
75
+
76
+ ## Versions and changelog
77
+ `v0.1.1`: first workable code, with function to generate a virtual cell. <br>
78
+ `v0.2.1`: added enumeration function <br>
79
+ `v0.2.2`: enumeration can be imported now (fix) <br>
80
+ `v0.3.0`: you can now make a batch of virtual cells<br>
81
+ `v0.4.3`: added tools to build a database
82
+
83
+ ## Debugging and support
84
+ The `virp` code has been tested on a limited number of platforms, so far Windows and Linux. If you are running into any problems during operation, please hound me (Andy Paul Chen) at la.vache.qui.vit(at)gmail.com, and I will try my best to help.
85
+
@@ -0,0 +1,14 @@
1
+ LICENSE
2
+ README.md
3
+ pyproject.toml
4
+ setup.py
5
+ virp/__init__.py
6
+ virp/database.py
7
+ virp/enumerate.py
8
+ virp/main.py
9
+ virp/matprop.py
10
+ virp.egg-info/PKG-INFO
11
+ virp.egg-info/SOURCES.txt
12
+ virp.egg-info/dependency_links.txt
13
+ virp.egg-info/requires.txt
14
+ virp.egg-info/top_level.txt
@@ -0,0 +1,5 @@
1
+ pymatgen
2
+ chgnet
3
+ matgl==1.0.0
4
+ dgl==1.1.2
5
+ poshcar
@@ -0,0 +1 @@
1
+ virp