We provided, designed for analyzing cybersecurity incidents, which is comprised of two primary task categories: understanding and generation, with a further breakdown into 28 subcategories of tasks.
The dataset is in question and answer format, using structured json format for understanding tasks and unstructured text format for generation tasks.
We also provide some multiple-choice questions to test the cognitive ability of the model in different vertical fields.… See the full description on the dataset page:
https://huggingface.co/datasets/Multilingual-Multimodal-NLP/SEVENLLM-Dataset.