A manually curated benchmark for evaluating chemistry and materials capabilities of Large Language Models
⚠️ IMPORTANT NOTICE - NOT FOR TRAINING
🚫 THIS DATASET IS STRICTLY FOR EVALUATION PURPOSES ONLY 🚫
DO NOT USE THIS DATASET FOR TRAINING OR FINE-TUNING MODELS
This benchmark is designed exclusively for evaluation and testing of existing models. Using this data for training would compromise the integrity of the benchmark and invalidate… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/ChemBench.