A ~100k diverse subset sampled from the molmo2 academic SFT mixture, meant to be mixed into another
training run as a replay set to prevent catastrophic forgetting of general language / image / video
abilities. It deliberately spans many tasks (per-dataset caps for diversity) and excludes long
videos (only short clips; videos > 10 MB filtered out) so it stays light on context length.… See the full description on the dataset page:
https://huggingface.co/datasets/weikaih/molmo2-replay-500k.