Project page | Paper | Code
🚨 Note: This repository now also contains the datasets for our the latest model in the Audio Flamingo series, Audio Flamingo Next.
AF-Chat is a high-quality fine-tuning dataset of ~75K multi-turn, multi-audio conversations (avg. 4.6 clips & 6.2 turns; range 2–8 clips & 2–10 turns) spanning speech, environmental sounds, and music. The dataset is partitioned into subsets based on each audio’s source dataset:… See the full description on the dataset page:
https://huggingface.co/datasets/nvidia/AF-Chat.