Evaluation benchmark for the Brief Review Agent of the Phytomni multi-agent system. The benchmark contains 71 complex plant-science topics, each paired with an expert-written reference report that serves as ground truth.
For each of the 71 topics, a reference report was written by a domain expert (PhD or PhD candidate) using the same knowledge corpus available to the AI agents. AI-generated reviews are scored against these expert references… See the full description on the dataset page:
https://huggingface.co/datasets/Phytomni/PhytoBench-Review.