Measuring general-assistant competence on real-world, multi-step, multi-modal tasks, scored by a verifiable exact-match reward.
Summary · Layout · Levels · Results · Analysis · Coverage · Dataset · Trajectories · Scoring · Reproduction · Verification
Terra measures whether an agent can complete real-world assistant tasks, not just answer trivia.
Each task is a question that is simple for a capable human but hard for… See the full description on the dataset page:
https://huggingface.co/datasets/ethara/terra-samples.