A broad collection of relational databases spanning many domains (academic,
e-commerce, finance, sports, biomedical, government, text2sql, and more), ported to the
RelBench manifest format. The Join is built for pretraining relational/tabular foundation
models: each database is self-describing and tasks ship labels as-is for large-scale
pretraining rather than held-out benchmarking.
Each dataset lives in its own subdirectory in the self-describing manifest layout (plain… See the full description on the dataset page:
https://huggingface.co/datasets/stanford-star/the-join.