With the rapid development of large language models (LLMs) in legal applications, systematically evaluating their reasoning ability in judgment prediction has become increasingly urgent. Currently, publicly available legal evaluation benchmarks lack a unified evaluation framework and do not adequately support these two tasks. To fill this gap, we propose MSLR, addressing a… See the full description on the dataset page:
https://huggingface.co/datasets/Yuwh07/MSLR-Bench.