STRABLE is a benchmarking corpus of 108 real-world tabular datasets containing string features,
designed to support empirical research on tabular machine learning pipelines that handle string entries.
Existing tabular benchmarks either exclude string columns or flatten them into fixed numerical
representations before evaluation, preventing the study of alternative string-handling… See the full description on the dataset page:
https://huggingface.co/datasets/strablebench/STRABLE.