Both manual transcripts and ASR outputs from the IWSLT2011 speech translation evalutation campaign are often used for the related punctuation annotation task. This dataset takes care of preprocessing said transcripts and automatically inserts punctuation marks given in the manual transcripts in the ASR outputs using Levenshtein aligment.