RLVR Dataset released by ASTRA: Automated Synthesis of agentic Trajectories and Reinforcement Arenas. The RLVR data is designed for training/evaluating tool use + multi-step reasoning with verifiable rewards in executable environments.
RLVR environments (Environment Synthesis): starting from QA pairs, we automatically decompose a main question into sub-questions and generate an executable tool environment (tool documentation /… See the full description on the dataset page:
https://huggingface.co/datasets/sunorme/astra_rlvr.