Rankings sourced from the EQ-Bench 3 canonical leaderboard data (2026-03-19 snapshot).
These are raw rubric scores, not the official ELO ranking — higher is higher but not
necessarily better (see eqbench.com for normalized ELO).
Newer models (gpt-5.4, claude-sonnet-4-6, claude-opus-4-6) are judged with Opus on the
live leaderboard and are not yet in the official repo data with Sonnet scores.
Qwen family comparison (all claude-3.7-sonnet judge)