This dataset dataset.jsonl consists of 10k samples of titles and conversations pairs in English.
The messaegs were assembled from Puffin and chatalpaca-20k. The titles were generated using gpt-3.5-turbo.
This dataset is part of a bigger project to fine-tune an LLM to generate short titles for chat conversations. You can find more information about it here:
https://github.com/ogrnz/generate-title-llm. The specific script used to generate the dataset is located here.… See the full description on the dataset page:
https://huggingface.co/datasets/ogrnz/chat-titles.