Chinese Simile (CS) Dataset
This dataset is constructed and based on the online free-access fictions that are tagged with sci-fi, urban novel, love story, youth, etc.
All similes are extracted by rich regular expression, and the extraction precision is estimated as 92% by labelling 500 random extracted samples. Further data filtering as well as processing is truly encouraged!
The data split in paper is as follows (You could find more… See the full description on the dataset page: https://huggingface.co/datasets/CreativeLang/wps_chinese_simile.