This is the training and validation query set used by the paper R2R: Efficiently Navigating Divergent Reasoning Paths with Small-Large Model Token Routing. This dataset contains token-level routing labels generated to train a lightweight router that selectively uses a Large Language Model (LLM) for critical, path-divergent tokens during inference, improving efficiency without sacrificing accuracy.
Roads to Rome (R2R) is a neural token router that efficiently combines Large Language Models… See the full description on the dataset page:
https://huggingface.co/datasets/nics-efc/R2R_query.