- Installation
- Dataset & checkpoints
- Training
- Inference
- Analysis
- CSP Pipeline
- Citation
- Acknowledgements
git clone https://github.com/Liu-Group-UF/MolCrystalFlow
cd MolCrystalFlowmamba env create -f environment.yml python=3.12
mamba activate molcrystalflow- All preprocessed datasets and model checkpoints are hosted on Zenodo: https://zenodo.org/records/19673190
- Dataset-specific preprocessing scripts are in
data-preprocess/.
The two preprocessed datasets can be downloaded as ZIP archives and extracted into data-preprocess/:
wget -P data-preprocess/ https://zenodo.org/record/19673190/files/thurlemann23.zip
wget -P data-preprocess/ https://zenodo.org/record/19673190/files/omc25-mcf.zip
unzip data-preprocess/thurlemann23.zip -d data-preprocess/
unzip data-preprocess/omc25-mcf.zip -d data-preprocess/After extracting, the preprocessed files will be available under:
data-preprocess/thurlemann23/preprocessed/— Thürlemann datasetdata-preprocess/omc25-mcf/preprocessed/— OMC25-MCF dataset
wget https://zenodo.org/record/19673190/files/model-checkpoints.zip
unzip model-checkpoints.zipAfter extracting, the checkpoints will be available under:
Pre-trained checkpoints with the lowest validation losses are provided in:
model-checkpoints/thurlemann23/- trained on the Thürlemann datasetmodel-checkpoints/omc25-mcf/- trained on the OMC25-MCF dataset
# Train on Thürlemann dataset
python molcrystalflow/experiments/train.py \
--config-name=molcrystal.yaml \
experiment.wandb.name=<experiment_name> \
experiment.trainer.max_epochs=300 \
data.cache_dir=./data-preprocess/thurlemann23/preprocessed/normalized
# Train on OMC25-MCF dataset
python molcrystalflow/experiments/train.py \
--config-name=omc25_molcrystal.yaml \
experiment.wandb.name=<experiment_name> \
experiment.trainer.max_epochs=500 \
model.bb_embedder.num_atom_types=12 \
data.cache_dir=./data-preprocess/omc25-mcf/preprocessed/normalizedTo continue training from <ckpt_path> in experiment <expname>
python molcrystalflow/experiments/train.py \
experiment.wandb.name=<expname> \
experiment.warm_start=<ckpt_path> \
+experiment.wandb.id=<run_id> \
+experiment.wandb.resume=must# Run inference with a trained checkpoint
python molcrystalflow/experiments/inference.py \
--config-name=inference.yaml \
inference.ckpt_path=<path/to/checkpoint.ckpt> \
data.cache_dir=./data-preprocess/thurlemann23/preprocessed/normalized \
interpolant.sampling.num_timesteps=50 \
interpolant.rots.exp_rate=3 \
interpolant.trans.scaling=9 \
inference.num_samples=10| Parameter | Description | Recommended |
|---|---|---|
interpolant.sampling.num_timesteps |
Number of integration steps | 50 |
interpolant.trans.scaling |
Scaling for centroid coordinates | 9.0 |
interpolant.rots.exp_rate |
Scaling for rotation orientations | 3.0 |
inference.num_samples |
Number of samples to generate per structure | 10 |
After inference completes, a predictions_K.pt file is generated in the inference/ subfolder of the checkpoint directory (where K is the number of samples). The following scripts analyze these predictions.
Evaluate generation qualtiy via pymatgen's StructureMatcher:
python molcrystalflow/experiments/run_structure_matching.py \
--pt_file <path/to/predictions_K.pt> \
--num_samples K \
--stol 0.8 \
--num_cpus 24This script:
- Generates ground truth (
gt_*.xyz) and predicted (pred_*.xyz) XYZ files - Performs structure matching using pymatgen StructureMatcher
- Saves RMSD results to CSV and matching summary to JSON
| Parameter | Description | Default |
|---|---|---|
--pt_file |
Path to predictions_K.pt file | Required |
--num_samples |
Number of samples per structure (K) |
Required |
--stol |
Atomic position tolerance | 0.8 |
--ltol |
Lattice length tolerance | 0.3 |
--angle_tol |
Lattice angle tolerance (degrees) | 10.0 |
--num_cpus |
CPUs for parallel processing | 24 |
--cg |
Coarse-grained matching | False |
Compare lattice volumes between ground truth and predicted structures:
python molcrystalflow/experiments/run_lattice_volume_analysis.py \
--gt_file <path/to/gt_*.xyz> \
--pred_file <path/to/pred_*.xyz> \
--num_samples K \
--output_dir <path/to/output>This script:
- Extracts lattice volumes from ground truth and predicted XYZ files
- Computes RMAD (Relative Mean Absolute Deviation) and per-structure deviations
- Generates publication-quality figures:
- KDE parity plot (PDF)
- Deviation boxplot (PDF)
- Lattice parameter comparison (PDF)
- Saves summary statistics to JSON
| Parameter | Description | Default |
|---|---|---|
--gt_file |
Path to ground truth XYZ file | Required |
--pred_file |
Path to predicted XYZ file | Required |
--num_samples |
Number of samples per structure | Required |
--output_dir |
Output directory for figures | Same as gt_file |
--prefix |
Prefix for output files | lattice_volume |
--kde_cmap |
Colormap for KDE parity plot | viridis |
# 1. Run inference
python molcrystalflow/experiments/inference.py \
--config-name=inference.yaml \
experiment.ckpt_path=./model-checkpoints/thurlemann23/best.ckpt \
data.cache_dir=./data-preprocess/thurlemann23/preprocessed/normalized \
inference.num_samples=10
# 2. Run structure matching (generates XYZ files)
python molcrystalflow/experiments/run_structure_matching.py \
--pt_file ./model-checkpoints/thurlemann23/inference/predictions_10.pt \
--num_samples 10 \
--stol 0.8
# 3. Run lattice volume analysis (with custom colormap)
python molcrystalflow/experiments/run_lattice_volume_analysis.py \
--gt_file ./model-checkpoints/thurlemann23/inference/gt_thurlemann23.xyz \
--pred_file ./model-checkpoints/thurlemann23/inference/pred_thurlemann23.xyz \
--num_samples 10 \
--kde_cmap plasmaNote:Crystal structure prediction pipeline is documented in the
csp-pipeline/folder.
If you find MolCrystalFlow or the processed open datasets useful, please cite:
@misc{zeng_molcrystalflow_2026,
title = {{MolCrystalFlow}: {Molecular} {Crystal} {Structure} {Prediction} via {Flow} {Matching}},
shorttitle = {{MolCrystalFlow}},
url = {http://arxiv.org/abs/2602.16020},
doi = {10.48550/arXiv.2602.16020},
urldate = {2026-02-26},
publisher = {arXiv},
author = {Zeng, Cheng and Sullivan, Harry W. and Egg, Thomas and Martirossyan, Maya M. and Höllmer, Philipp and Jin, Jirui and Hennig, Richard G. and Roitberg, Adrian and Martiniani, Stefano and Tadmor, Ellad B. and Liu, Mingjie},
month = feb,
year = {2026},
note = {arXiv:2602.16020 [cs]},
keywords = {Computer Science - Machine Learning, Condensed Matter - Materials Science},
}If you use the Thürlemann dataset, please also cite:
@article{thurlemann_regularized_2023,
title = {Regularized by {Physics}: {Graph} {Neural} {Network} {Parametrized} {Potentials} for the {Description} of {Intermolecular} {Interactions}},
volume = {19},
issn = {1549-9618},
shorttitle = {Regularized by {Physics}},
url = {https://doi.org/10.1021/acs.jctc.2c00661},
doi = {10.1021/acs.jctc.2c00661},
number = {2},
urldate = {2026-02-26},
journal = {Journal of Chemical Theory and Computation},
publisher = {American Chemical Society},
author = {Thürlemann, Moritz and Böselt, Lennard and Riniker, Sereina},
month = jan,
year = {2023},
pages = {562--579},
}License: The Thürlemann dataset is licensed under the CC BY-NC-SA 4.0 license (Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International). Original raw HDF data and the corresponding
README.mdcan be retrieved from the link.
If you use the OMC25-MCF dataset, please cite:
@misc{gharakhanyan2025_OMC25,
title = {Open {Molecular} {Crystals} 2025 ({OMC25}) {Dataset} and {Models}},
url = {http://arxiv.org/abs/2508.02651},
doi = {10.48550/arXiv.2508.02651},
urldate = {2025-08-05},
publisher = {arXiv},
author = {Gharakhanyan, Vahe and Barroso-Luque, Luis and Yang, Yi and Shuaibi, Muhammed and Michel, Kyle and Levine, Daniel S. and Dzamba, Misko and Fu, Xiang and Gao, Meng and Liu, Xingyu and Ni, Haoran and Noori, Keian and Wood, Brandon M. and Uyttendaele, Matt and Boromand, Arman and Zitnick, C. Lawrence and Marom, Noa and Ulissi, Zachary W. and Sriram, Anuroop},
month = aug,
year = {2025},
keywords = {Physics - Chemical Physics},
}License: The OMC25 dataset is provided under a CC BY 4.0 license (Creative Commons Attribution 4.0 International). Original raw ase-db OMC25 data can be found at huggingface
MolCrystalFlow builds upon the following projects:
