A unified evaluation dataset aggregating multiple legal reasoning benchmarks into a single flat schema for cost-efficient LLM evaluation. All samples are pre-formatted as zero-shot prompts — ready to send directly to a model.
Total: ~9,769 samples across 202 tasks from 5 source benchmarks.
benchmark
Source benchmark: legalbench, barexam, lexam, housingqa, or legal_hallucinations
task_name… See the full description on the dataset page:
https://huggingface.co/datasets/Pq234/legal-eval.