It's 40,000 rows sampled from teknium/openhermes (not the newer 2.5).
Filtered some GPTisms I dislike out, and removed rows with short output as well to bias towards longer answers.
bad_phrases = ["couldn't help but", "can't resist", "random", "unethical", "I'm sorry, but", "I'm sorry but", "as an AI", "as a Language Model", "AI Language Model", "language model", "However, it is important to", "However, it's important", "ethical guidelines", "just an AI", "within my programming", "illegal"… See the full description on the dataset page:
https://huggingface.co/datasets/lodrick-the-lafted/Hermes-40K.