The tokenized instruction-tuning dataset used for the supervised fine-tuning (SFT)
campaign of pinkelephantlimited/pinkelephant-llm-48b-s.
Source: teknium/OpenHermes-2.5 (train split, 1,001,551 conversations)
Tokenization: Pink Elephant 48B-S tokenizer (ChatML template from the model repo,
chat_template.jinja), sequence truncated to 1,024 tokens
Columns: input_ids / labels — raw… See the full description on the dataset page:
https://huggingface.co/datasets/pinkelephantlimited/pinkelephant-llm-48b-sft-openhermes-1m.