MLCR is a long-context evaluation benchmark designed to assess how effectively large language models (LLMs) answer questions based on medical and insurance documents. The benchmark is built from 10 synthetic cases that closely resemble real-world patient and claims scenarios, testing model performance as relevant information becomes progressively diluted with unrelated content. It contains 205 questions distributed across… See the full description on the dataset page:
https://huggingface.co/datasets/Wisedocs/mlcr-dataset.