Dataset Card for Estonian Parliament summary dataset
Dataset Summary
This dataset is collected from Estonian parliament stenograms. Each text contains speaker and his/her text (and could possibly contain multiple speaker texts).
Texts are concatenated so that maximum number of tokens would not exceed 2048 (for mBart and similar models).
This dataset is created to train a bit longer text summaries than default transformer models allow.
Supported Tasks… See the full description on the dataset page: https://huggingface.co/datasets/rristo/et_parliament_stenos_summary.