This is the Dataset of Degarbayan-SC paper.
You can Finetune with this dataset on the transformers models using Github.
from datasets import load_dataset
dataset = load_dataset("m0javad/Degarbayan-SC-dataset")
our sentence length distribution is between 3 and 19 words and sentences are an average of 8 words. This makes sense because in the movie subtitles, sentences are shown in a range of… See the full description on the dataset page:
https://huggingface.co/datasets/m0javad/Degarbayan-SC-dataset.