This is a record of the dataset collection used to train the huginn-0125 model. The data is provided in a semi-prepared format. We provide 4096 parquet files for train and val each which contain the exact rows used for training and validation (on the 4096 accelerators the model was trained on).
Each row is 4097 tokens long, which includes formatting tokens. The tokenizer here is the same as the model,
https://huggingface.co/tomg-group-umd/huginn-0125. However… See the full description on the dataset page:
https://huggingface.co/datasets/tomg-group-umd/huginn-dataset.