To effectively measure the alignment between automatic evaluation metrics and radiologists' assessments in medical text generation tasks,
we have established a comprehensive benchmark, RaTE-Eval,
that encompasses three tasks, each with its official test set for fair comparison, as detailed below.
The comparison of RaTE-Eval Benchmark and existed radiology report evaluation Benchmark is listed in Table.… See the full description on the dataset page:
https://huggingface.co/datasets/Angelakeke/RaTE-Eval.