This repository contains a range of Arena-Hard-Auto benchmark artifacts sourced as part of the 2024 paper Style Outweighs Substance.
Repository Structure
Model Responses for Arena Hard Auto Questions: data/ArenaHardAuto/model_answer
Our standard reference model for pairwise comparisons was gpt-4-0314.
Llama-3-8B Variants: bagel-8b-v1.0, Llama-3-8B-Magpie-Align-SFT-v0.2, Llama-3-8B-Magpie-Align-v0.2, Llama-3-8B-Tulu-330K, Llama-3-8B-WildChat… See the full description on the dataset page:
https://huggingface.co/datasets/nyu-dice-lab/sos-artifacts.