Ultraset - all-in-one dataset for SFT training in Alpaca format
About the dataset
This dataset is designed to facilitate training and retraining of LLM models using the SFT method in the Alpaca format.
Brief information
Number of rows: 785K
Type of dataset files: parquet
Type of dataset: text, alpaca
Languages:
English
Russian
French
Italian
Spanish
German
Chinese
Korean
License: flexible multi-license, main - MIT