Prompts broadcast across every model in a serving pool, with per-model responses and
reference-grounded judge scores. Each record gives directly comparable outputs for the same
input, which is what makes the corpus usable for routing research — learning which model in
a heterogeneous pool will answer a given prompt well — rather than only for single-model
evaluation.
Two model families are provided. They share one prompt taxonomy and one… See the full description on the dataset page:
https://huggingface.co/datasets/asdwb/route_balance_model_estimator.