This dataset combines all original CEFR-level sentences from training, validation, and test sets (preserving all paid annotator data) with synthetic A2 and C1 sentences generated by a fine-tuned LLaMA-3-8B model. Synthetic sentences were validated using a fine-tuned MLP classifier (~93% accuracy) to ensure the predicted CEFR level is within 1 level of the target (e.g., A2 accepts A1, A2, B1; C1 accepts B2, C1, C2). Duplicate sentences were… See the full description on the dataset page:
https://huggingface.co/datasets/Mr-FineTuner/CEFR_Mixed_Dataset_A2_C1_1.