This dataset consists of multiple samples, each containing the following fields:
tag: The label of the sample, indicating the category of the content, such as "pornographic".
title: The title of the video, usually containing the username and user ID.
OCR: Optical Character Recognition results, extracted text from images.
ASR: Automatic Speech Recognition results, extracted text from audio.
images: A list of image filenames, representing… See the full description on the dataset page: https://huggingface.co/datasets/RedZhu/KuaiMod.