This dataset includes text data that has been annotated for rewriting using the manual annotation method described in our paper to enhance the performance of speech translation systems. It is based on the original BSTC dataset.
Dataset Structure
The dataset consists of pairs of segments structured as follows:
left: The ASR transcript.
right: The human annotation text.