A unified collection of 5 high-quality question-answering and reasoning datasets in VERL format, deduplicated and optimized for reinforcement learning training.
This dataset combines 5 diverse QA and reasoning datasets into a single unified collection:
Total Problems: 86,379 unique problems (after 0.00% deduplication)
Original Size: 0 problems (before deduplication)
Format: VERL (Volcano Engine Reinforcement Learning)
Language:… See the full description on the dataset page:
https://huggingface.co/datasets/sungyub/qa-verl-unified.