This dataset contains the MMHS150K samples that were not used in the 3,000-sample manual annotation set.
Rows: 148,623
Columns:
input: ordered list of content blocks. Each block has type and content. MMHS150K rows contain tweet text followed by an image file reference.
original_label: original MMHS150K labels_str annotator labels.
Media files are stored in multimodal_files.zip; image content values in input correspond to filenames inside that zip.