⚠️ Part of the BidirLM-Omni Collection > This dataset is a specific modality sub-sample of the corpus used to train the BidirLM-Omni models.
Looking for the full training mixture? > If you want to access the complete, balanced 1.8M sample omnimodal dataset (integrating text, image, audio), please visit the global integration hub here:👉 BidirLM/BidirLM-Omni-Contrastive
If you use this processed dataset or the broader BidirLM mixture in your research, please cite… See the full description on the dataset page:
https://huggingface.co/datasets/BidirLM/laion_audio_contrastive.