A derived redistribution, not new data. These are the datasets used by the
OPD / TRD pipeline, already converted into the format the pipeline reads, so a
target machine with a poor network can skip the download-and-rebuild step
entirely.
Four training domains (math, code, instruct, stem) and seven math eval sets.
~4.3 GB total, of which TACO is ~4 GB; everything else is under 250 MB.
Path
Rows
Source repo… See the full description on the dataset page:
https://huggingface.co/datasets/HzChen20/wheel-opd-data.