A large-scale synthetically generated paraphrase dataset containing sentence pairs with balanced positive and negative examples across varied domains and writing styles.
Dataset Details
Dataset Description
Name: llm-paraphrases
Summary: A synthetic paraphrase dataset generated using large language models, designed for training embedding models for semantic caching and paraphrase detection. Each example contains a pair of… See the full description on the dataset page: https://huggingface.co/datasets/redis/llm-paraphrases.