AgReason is an expert-curated benchmark designed to evaluate large language models (LLMs) on complex, contextual agricultural reasoning. It contains 100 open-ended questions, each paired with gold-standard answers created and reviewed by agronomy experts. These questions are derived from real-world farming scenarios and require multi-step reasoning over location-specific, seasonal, and environmental constraints.
🧠 Benchmark Overview… See the full description on the dataset page: https://huggingface.co/datasets/BGLab/AgReason.