A 1,200-sample curated benchmark dataset for evaluating LLMs on factual accuracy and truthfulness.
Sourced from MMLU and TruthfulQA, cleaned and formatted for the
LLM Evaluation Framework.
Dataset Summary
Split
Samples
Use
train
500
Fine-tuning reference / training baselines
validation
200
Hyperparameter tuning
test
500
Final benchmark — use this for fair comparisons
Total
1,200
Quick… See the full description on the dataset page: https://huggingface.co/datasets/vigneshwar234/llm-eval-benchmark.