PubChemQCR dataset contains the relaxation trajectory of ~3.5 million small molecules, which can facilitate the development of machine learning interatomic potential (MLIP) models. The relaxation is performed sequentially using PM3, Hartree-Fock, and DFT methods, resulting in a total of 300 million snapshots, 105 million of which are computed using DFT. The dataset is split into two portions, a subset and a full set. Both sets share the same… See the full description on the dataset page:
https://huggingface.co/datasets/divelab/PubChemQCR.