Compare DeepSearch on Rag, vannilla models on complex dataset.
This dataset contains 6 experiments
from the EvalAP evaluation platform.
Datasets: WikipediaFrames_150
Metrics: answer_relevancy, judge_exactness, judge_notator, output_length
Scores
WikipediaFrames_150
deepsearch_8B(3.1)70B(3.3)-web_3_3_3
0.73 ± 0.44
0.27 ± 0.44
3.62 ±… See the full description on the dataset page:
https://huggingface.co/datasets/AgentPublic/evalap-wikipedia_frames_150-11.