FSS1 (Fast Start Set vol.1) is a practical pretraining dataset for English and Russian causal language models.
This thing was built for a very specific use case: you want a model that can start speaking, reasoning, continuing text, and handling dialogue without paying the full price of classic large-scale web pretraining. So this is not a "pure SFT set", not a sterile benchmark soup, and not a… See the full description on the dataset page:
https://huggingface.co/datasets/srs6901/FSS1.