This dataset contains 8,230 table reasoning samples from 3 datasets (HiTab, MultiHierTT, FinQA) for reinforcement learning training with VERL (Volcano Engine Reinforcement Learning). The data is extracted and preprocessed from LLM360/guru-RL-92k.
Guru is a reasoning model trained using cross-domain reinforcement learning. This dataset focuses on table reasoning tasks where models must analyze hierarchical tables and financial data to answer… See the full description on the dataset page:
https://huggingface.co/datasets/sungyub/guru-table-verl.