Skip to content

Latest commit

 

History

37 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

MolCrystalFlow: Molecular Crystal Structure Prediction with Flow Matching

arXiv

Table of Contents


Installation

Clone repository

git clone https://github.com/Liu-Group-UF/MolCrystalFlow 
cd MolCrystalFlow

Setup environment

mamba env create -f environment.yml python=3.12
mamba activate molcrystalflow

Dataset & checkpoints

Download preprocessed datasets

The two preprocessed datasets can be downloaded as ZIP archives and extracted into data-preprocess/:

wget -P data-preprocess/ https://zenodo.org/record/19673190/files/thurlemann23.zip
wget -P data-preprocess/ https://zenodo.org/record/19673190/files/omc25-mcf.zip
unzip data-preprocess/thurlemann23.zip -d data-preprocess/
unzip data-preprocess/omc25-mcf.zip -d data-preprocess/

After extracting, the preprocessed files will be available under:

  • data-preprocess/thurlemann23/preprocessed/ — Thürlemann dataset
  • data-preprocess/omc25-mcf/preprocessed/ — OMC25-MCF dataset

Download model checkpoints

wget https://zenodo.org/record/19673190/files/model-checkpoints.zip
unzip model-checkpoints.zip

After extracting, the checkpoints will be available under:

Pre-trained checkpoints with the lowest validation losses are provided in:

  • model-checkpoints/thurlemann23/ - trained on the Thürlemann dataset
  • model-checkpoints/omc25-mcf/ - trained on the OMC25-MCF dataset

Training

Quick Start

# Train on Thürlemann dataset
python molcrystalflow/experiments/train.py \
    --config-name=molcrystal.yaml \
    experiment.wandb.name=<experiment_name> \
    experiment.trainer.max_epochs=300 \
    data.cache_dir=./data-preprocess/thurlemann23/preprocessed/normalized

# Train on OMC25-MCF dataset
python molcrystalflow/experiments/train.py \
    --config-name=omc25_molcrystal.yaml \
    experiment.wandb.name=<experiment_name> \
    experiment.trainer.max_epochs=500 \
	model.bb_embedder.num_atom_types=12 \ 
    data.cache_dir=./data-preprocess/omc25-mcf/preprocessed/normalized

Resume training from checkpoint

To continue training from <ckpt_path> in experiment <expname>

python molcrystalflow/experiments/train.py \
    experiment.wandb.name=<expname> \
    experiment.warm_start=<ckpt_path> \
    +experiment.wandb.id=<run_id> \
    +experiment.wandb.resume=must

Inference

Quick Start

# Run inference with a trained checkpoint
python molcrystalflow/experiments/inference.py \
    --config-name=inference.yaml \
    inference.ckpt_path=<path/to/checkpoint.ckpt> \
    data.cache_dir=./data-preprocess/thurlemann23/preprocessed/normalized \
    interpolant.sampling.num_timesteps=50 \
    interpolant.rots.exp_rate=3 \
    interpolant.trans.scaling=9 \
    inference.num_samples=10

Key Parameters

Parameter Description Recommended
interpolant.sampling.num_timesteps Number of integration steps 50
interpolant.trans.scaling Scaling for centroid coordinates 9.0
interpolant.rots.exp_rate Scaling for rotation orientations 3.0
inference.num_samples Number of samples to generate per structure 10

Analysis

After inference completes, a predictions_K.pt file is generated in the inference/ subfolder of the checkpoint directory (where K is the number of samples). The following scripts analyze these predictions.

Structure Matching

Evaluate generation qualtiy via pymatgen's StructureMatcher:

python molcrystalflow/experiments/run_structure_matching.py \
    --pt_file <path/to/predictions_K.pt> \
    --num_samples K \
    --stol 0.8 \
    --num_cpus 24

This script:

  1. Generates ground truth (gt_*.xyz) and predicted (pred_*.xyz) XYZ files
  2. Performs structure matching using pymatgen StructureMatcher
  3. Saves RMSD results to CSV and matching summary to JSON
Parameter Description Default
--pt_file Path to predictions_K.pt file Required
--num_samples Number of samples per structure (K) Required
--stol Atomic position tolerance 0.8
--ltol Lattice length tolerance 0.3
--angle_tol Lattice angle tolerance (degrees) 10.0
--num_cpus CPUs for parallel processing 24
--cg Coarse-grained matching False

Lattice Volume Comparison

Compare lattice volumes between ground truth and predicted structures:

python molcrystalflow/experiments/run_lattice_volume_analysis.py \
    --gt_file <path/to/gt_*.xyz> \
    --pred_file <path/to/pred_*.xyz> \
    --num_samples K \
    --output_dir <path/to/output>

This script:

  1. Extracts lattice volumes from ground truth and predicted XYZ files
  2. Computes RMAD (Relative Mean Absolute Deviation) and per-structure deviations
  3. Generates publication-quality figures:
    • KDE parity plot (PDF)
    • Deviation boxplot (PDF)
    • Lattice parameter comparison (PDF)
  4. Saves summary statistics to JSON
Parameter Description Default
--gt_file Path to ground truth XYZ file Required
--pred_file Path to predicted XYZ file Required
--num_samples Number of samples per structure Required
--output_dir Output directory for figures Same as gt_file
--prefix Prefix for output files lattice_volume
--kde_cmap Colormap for KDE parity plot viridis

Example: Full Analysis Pipeline

# 1. Run inference
python molcrystalflow/experiments/inference.py \
    --config-name=inference.yaml \
    experiment.ckpt_path=./model-checkpoints/thurlemann23/best.ckpt \
    data.cache_dir=./data-preprocess/thurlemann23/preprocessed/normalized \
    inference.num_samples=10

# 2. Run structure matching (generates XYZ files)
python molcrystalflow/experiments/run_structure_matching.py \
    --pt_file ./model-checkpoints/thurlemann23/inference/predictions_10.pt \
    --num_samples 10 \
    --stol 0.8

# 3. Run lattice volume analysis (with custom colormap)
python molcrystalflow/experiments/run_lattice_volume_analysis.py \
    --gt_file ./model-checkpoints/thurlemann23/inference/gt_thurlemann23.xyz \
    --pred_file ./model-checkpoints/thurlemann23/inference/pred_thurlemann23.xyz \
    --num_samples 10 \
    --kde_cmap plasma

CSP Pipeline

Note:Crystal structure prediction pipeline is documented in the csp-pipeline/ folder.

Citation

If you find MolCrystalFlow or the processed open datasets useful, please cite:

@misc{zeng_molcrystalflow_2026,
	title = {{MolCrystalFlow}: {Molecular} {Crystal} {Structure} {Prediction} via {Flow} {Matching}},
	shorttitle = {{MolCrystalFlow}},
	url = {http://arxiv.org/abs/2602.16020},
	doi = {10.48550/arXiv.2602.16020},
	urldate = {2026-02-26},
	publisher = {arXiv},
	author = {Zeng, Cheng and Sullivan, Harry W. and Egg, Thomas and Martirossyan, Maya M. and Höllmer, Philipp and Jin, Jirui and Hennig, Richard G. and Roitberg, Adrian and Martiniani, Stefano and Tadmor, Ellad B. and Liu, Mingjie},
	month = feb,
	year = {2026},
	note = {arXiv:2602.16020 [cs]},
	keywords = {Computer Science - Machine Learning, Condensed Matter - Materials Science},
}

If you use the Thürlemann dataset, please also cite:

@article{thurlemann_regularized_2023,
	title = {Regularized by {Physics}: {Graph} {Neural} {Network} {Parametrized} {Potentials} for the {Description} of {Intermolecular} {Interactions}},
	volume = {19},
	issn = {1549-9618},
	shorttitle = {Regularized by {Physics}},
	url = {https://doi.org/10.1021/acs.jctc.2c00661},
	doi = {10.1021/acs.jctc.2c00661},
	number = {2},
	urldate = {2026-02-26},
	journal = {Journal of Chemical Theory and Computation},
	publisher = {American Chemical Society},
	author = {Thürlemann, Moritz and Böselt, Lennard and Riniker, Sereina},
	month = jan,
	year = {2023},
	pages = {562--579},
}

License: The Thürlemann dataset is licensed under the CC BY-NC-SA 4.0 license (Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International). Original raw HDF data and the corresponding README.md can be retrieved from the link.

If you use the OMC25-MCF dataset, please cite:

@misc{gharakhanyan2025_OMC25,
	title = {Open {Molecular} {Crystals} 2025 ({OMC25}) {Dataset} and {Models}},
	url = {http://arxiv.org/abs/2508.02651},
	doi = {10.48550/arXiv.2508.02651},
	urldate = {2025-08-05},
	publisher = {arXiv},
	author = {Gharakhanyan, Vahe and Barroso-Luque, Luis and Yang, Yi and Shuaibi, Muhammed and Michel, Kyle and Levine, Daniel S. and Dzamba, Misko and Fu, Xiang and Gao, Meng and Liu, Xingyu and Ni, Haoran and Noori, Keian and Wood, Brandon M. and Uyttendaele, Matt and Boromand, Arman and Zitnick, C. Lawrence and Marom, Noa and Ulissi, Zachary W. and Sriram, Anuroop},
	month = aug,
	year = {2025},
	keywords = {Physics - Chemical Physics},
}

License: The OMC25 dataset is provided under a CC BY 4.0 license (Creative Commons Attribution 4.0 International). Original raw ase-db OMC25 data can be found at huggingface

Acknowledgements

MolCrystalFlow builds upon the following projects:

Releases

Packages

Contributors

Languages