TerraCoT is a remote-sensing chain-of-thought VQA dataset in which every [SEG] token in
an answer is grounded by a pixel-level ground-truth mask. Each sample pairs an optical
image with a reasoning trace, and (where applicable) a second modality in image2
(Sentinel-1 SAR, or a bi-temporal companion image).
Two splits:
train — samples that carry an ordered labels list aligned to the [SEG] tokens.
vqa — question/answer samples without the labels column.… See the full description on the dataset page: https://huggingface.co/datasets/sy1998/TerraCoT.