RLVR-CoS (Reinforcement Learning with Verifiable Rewards — Chain-of-Sight) is a visual reasoning dataset designed to train and evaluate vision-language models on structured, spatially-grounded image understanding. Each sample pairs an image with a Chain-of-Sight (CoS) reasoning trace that decomposes free-form visual reasoning into atomic, verifiable steps.