Balanced Bilingual Literary Complexity Dataset (EN + ZH)
Description
This dataset contains paragraph-level literary segments labeled with CEFR-style difficulty levels (A/B/C).
Languages:
text: paragraph
difficulty: A/B/C
language: en/zh
open_source_books: original source
English: CommonLit readability model,syntactic complexity, the proportion of rare… See the full description on the dataset page:
https://huggingface.co/datasets/merriamtao/BiLit-CEFR.