Evaluation benchmark for the Analyst Agent of the Phytomni multi-agent system. The benchmark covers 10 bioinformatics domains with five tasks per domain, totalling 50 end-to-end analysis scenarios. Each task ships with a binary attachment containing the input dataset(s) the agent needs in order to plan and execute the analysis.
For each of the 10 domains, five datasets of similar task type were assembled. The 10 domains are:… See the full description on the dataset page:
https://huggingface.co/datasets/Phytomni/PhytoBench-Analysis.