This dataset contains 8 experiments
from the EvalAP evaluation platform.
Datasets: MFS_questions_v01
Models evaluated: meta-llama/Llama-4-Scout-17B-16E-Instruct
Metrics: answer_relevancy, generation_time, judge_exactness, judge_notator, output_length