This dataset is a curated collection of instruction-style samples designed for training and fine-tuning large language models. Each example consists of an input and a corresponding output, forming a structured interaction suitable for supervised learning.
The dataset has been processed and organized based on token length, enabling efficient training across different context sizes.
📊 Dataset Splits… See the full description on the dataset page: https://huggingface.co/datasets/epoch-lab/codex-m-10k.