The Tremendous TabLib Trawl (T4) is a dataset for training tabular foundation models.
The dataset is described in detail in our paper, "Large Scale Transfer Learning for Tabular Data via Language Modeling."
The paper also includes a datasheet for this dataset.
T4 consists of a set of Parquet files (described below). For examples and infrastructure showing how to train a lannguage model
on T4, see our open-source Python library, rtfm, which was used to train TabuLa-8B on T4.
Files… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/t4-full.