synCUB is a synthetic, paired-image benchmark for evaluating concept-based
interpretability. Each item is an (original, synthetic) image pair that differs
in exactly one CUB attribute: the original contains old_attr, and the
synthetic image replaces it with new_attr (e.g. has_breast_pattern::solid →
has_breast_pattern::spotted). Images are generated with FLUX.2 [dev]
conditioned on CUB reference images.
It accompanies the paper "Evaluating the Interpretability of Sparse… See the full description on the dataset page:
https://huggingface.co/datasets/jokl/syncub.