I(q)@L=50.h5
~66 GB
HDF5 database of I(q) curves and molecular data
iq_train_set-ENCODING.sqlite3
~860 MB
Encoding index: maps every molecule to its atom count and VOCAB indices, so the data pipeline never needs to scan the 66 GB HDF5 file… See the full description on the dataset page: https://huggingface.co/datasets/noshou/iq_train_set.