The Lit2Vec Subfield Classifier Dataset is a curated and preprocessed collection of scientific research metadata designed for text classification and embedding-based machine learning tasks.It includes over 39,900 chemistry abstract and tldr text annotated with domain subfields, dense text embeddings, and structured metadata, making it suitable for: