Multimodal molecular dataset used to train Mol-JEPA (a multimodal Joint
Embedding Predictive Architecture for molecules). Each row of metadata.csv
describes one molecule (SMILES + InChIKey + source dataset + labels) and points to
precomputed per-modality embedding/target files stored as NumPy arrays.
These are the modalities included (note that not every modality is available for every row - there is quite some sparsity). For detailed… See the full description on the dataset page:
https://huggingface.co/datasets/Flogrammer/Mol-JEPA-dataset.