This dataset contains text extracted from archived Orkut community pages. It is derived from Internet Archive / Archive Team Orkut WARC and CDX files and converted into a hierarchical Parquet dataset for research use.
The dataset focuses on human-authored text fields found in community and forum pages, including community names and descriptions, forum topic metadata, usernames, reply dates, reply titles, and reply bodies.… See the full description on the dataset page: https://huggingface.co/datasets/SalatielJordao/orkut-communities.