š CrossEval: Benchmarking LLM Cross Capabilities
Release of Model Responses and Evaluations
In addition to the CrossEval benchmark, we release the responses from all 17 models in this repository, along with the ratings and explanations provided by GPT-4o as the evaluator. The included model families are:
GPT
Claude
Gemini
Reka
Dataset Structure
Each instance in the dataset contains the following fields: