This is a small dataset I've manually curated and annotated designed to make higher quality, more readable transcripts.
It's used to finetune my conciser model. The dataset currently consists of ~50 examples of paragraphs of
text taken from transcripts of various podcasts, and then lightly touchced up to enhance readability.
Some examples of edits I make involve
removing filler words
breaking up and rearranging long run-on sentences… See the full description on the dataset page:
https://huggingface.co/datasets/chrislee973/llama3-conciser-dataset.