This is a filtered combination of three instruction-tuning datasets (in LLAMA chat format) synthetically derived from Arxiv abstracts. These are:
ArtifactAI/arxiv-math-instruct-50k
ArtifactAI/arxiv-cs-ml-instruct-tune-50k
ArtifactAI/arxiv-physics-instruct-tune-30k
The final dataset is derived by running this concatenated mix via a KenLM model (OSCAR EN model at edugp\kenlm) with a perplexity filter of 400.