LV-Eval, a bilingual benchmark dataset targeted to evaluate long context large language models with fairer tasks and metrics. Our benchmark includes 12 finegrained tasks and each task is composed of 5 length levels of 16k, 32k, 64k, 128k, 256k, respectively, with balanced amount of questions.