This dataset has been created using the transcripts from lexicap.
Each transcript has been partitioned into chunks of max 1000 tokens.
GPT-3.5 has been used to augment the chunks with a description and context field.
The features provided are: title, description, context, transcript.
The LexiGPT-Podcast-Corpus dataset offers a comprehensive collection of transcripts from the Lex Fridman podcast, thoughtfully curated and… See the full description on the dataset page:
https://huggingface.co/datasets/Philipp-Sc/LexiGPT-Podcast-Corpus.