This dataset represents the publicly available collection of data scraped from the virgool.io website. The data extraction was strategically performed based on specific tags and user. The dataset comprises approximately 62,000 entries across several key attributes: title, text, tags, likes, replies, reading_time, user_id, and URL.
This resource is particularly beneficial for researchers and developers aiming to pre-train large language models (LLMs), as the 'text' column provides a rich corpus… See the full description on the dataset page:
https://huggingface.co/datasets/Msobhi/virgool_62k.