MiroEval: Benchmarking Multimodal Deep Research Agents in Process and Outcome
MiroEval is a comprehensive evaluation framework for Deep Research systems, providing automated task generation and assessment across three complementary dimensions: Factual correctness, Point-wise quality, and Process quality.
All three evaluation modules share a single Python environment managed by uv at the repo root:
uv sync
If you use… See the full description on the dataset page:
https://huggingface.co/datasets/miromind-ai/MiroEval-data.