Given a user prompt, the model generates a structured evaluation rubric in [Hard Rule] / [Principle] format. These rubrics are used to judge LLM response quality.
Evaluation
~83.5% format validity on Chatbot Arena prompts
Used as the baseline rubric generator in the GRUBRIC pipeline