A travel dataset with image audio modality, stored in npy sharded format.
Preprocessing & Augmentation
Preprocessing: curriculum
Augmentation: none
Split strategy: random 90 10
Sampling: weighted
Quality filtering: strict
Labeling: semi auto
clean.py — main artifact of this repository
See the license field… See the full description on the dataset page:
https://huggingface.co/datasets/sebastianwalter/translate-lite.