This dataset combines all original CEFR-level sentences from training, validation, and test sets (preserving all paid annotator data) with synthetic C2-level sentences generated by a fine-tuned LLaMA-3-8B model. Synthetic sentences were validated using a fine-tuned MLP classifier (~93% accuracy) to ensure the predicted CEFR level is within 1 level of C2 (i.e., C1 or C2). Duplicate sentences were rejected to ensure diversity. Synthetic data was… See the full description on the dataset page:
https://huggingface.co/datasets/Mr-FineTuner/CEFR_Mixed_Dataset_C2_1.