The books dataset consists of a diverse collection of books organized into 9 categories, it splitted to train, validation where the train contains 40 books, and the validation 9 books.
This dataset is cleaned well and designed to support various natural language processing (NLP) tasks, including text generation and masked language modeling.
title: The tilte of the book.
author: The author of the book.
category: The… See the full description on the dataset page:
https://huggingface.co/datasets/DhruvExploring/books.