For more information, refer to our blogpost. We used these datasets for long instruction-following training. The maximal sequence length of the examples is 32,768.
Synthetic-ConvQA with RAFT-style augmentation.
Our synthetic long-context data is based on an approach introduced by [Zhang et al., 2024] called Retrieval Augmented Fine-Tuning (RAFT). For each… See the full description on the dataset page:
https://huggingface.co/datasets/cerebras/Synth-Long-SFT32K.