Data for training molvae — a SELFIES molecular VAE (matryoshka latent) for Bayesian
optimization of mechanophores. Two layers:
raw/ — the source chemical databases (SMILES/SELFIES, pre-tokenized), as collected.
mixes/ — the derived training datasets: weighted, shuffled, materialized token-shard
"mixes" built from the raw sources. Each mix is documented by its config + README + build
code, so it is fully reproducible. The multi-GB token shards themselves live… See the full description on the dataset page:
https://huggingface.co/datasets/MechanophoresResearch/AutoencoderDataset.