Procedural rows from reasoning-gym, formatted for verl GRPO.
{
"config": "/home/owais/Projects/rlvr/rlvr/configs/datasets/cryptarithm-instruct copy.yaml",
"template_type": "qwen-instruct",
"developer_prompt": null,
"data_source": "reasoning_gym",
"default_extract": "answer_tag",
"train_rows": 100000,
"test_rows": 4096,
"train_seed": 42,
"test_seed": 43,
"tasks": {
"cryptarithm": {
"weight": 1… See the full description on the dataset page:
https://huggingface.co/datasets/carbonteq/rg-cryptarithm-instruct-100k.