This is an LLM rated version of euclaise/reddit-instruct-curated, which is already a good dataset imo.
Only post titles and comment texts were rated as post texts can be confusing due to edits and seemingly out of context information.
First, I filtered examples with <250 comment score. Of course this is not a very efficient filtering as some pairs might have references to other comments or simply be unhelpful, yet upvoted due to Reddit hivemind.
Next I sent the example pairs with a rating… See the full description on the dataset page:
https://huggingface.co/datasets/Ba2han/Reddit-instruct-curated_rated-1.2k.