The instruction-tuning corpus behind Pink Elephant 48B — pre-tokenized and ready to train.
This dataset contains 99,661 training examples and 1,000 evaluation examples used to fine-tune the Pink Elephant 48B MoE model. All examples are already tokenized with the Pink Elephant tokenizer (vocabulary 100,352), so training can begin with no preprocessing.