This evaluation is conducted on the advent of code dataset on several models including Qwen2.5-Coder-32B-Instruct, DeepSeek-V3-fp8, Llama-3.3-70B-Instruct, GPT-4o-mini, DeepSeek-R1.The aim is to to see how well these models can handle real-world puzzle prompts, generate correct Python code, and ultimately shed light on which LLM truly excels at reasoning and problem-solving.We used pass@1 to measure the functional correctness.
Results… See the full description on the dataset page: https://huggingface.co/datasets/Chemin-AI/advent_of_code_evaluations.