ArtifactsBench: Bridging the Visual-Interactive Gap in LLM Code Generation Evaluation
Tencent Hunyuan Team
📖 Paper •
🏠 Home Page •
💻 Code •
🏆 Leaderboard •
📜 Citation
Figure 1: Automation level versus human–alignment across evaluation frameworks. The red star marks the fully manual WebDev Arena (100% human effort), while the blue bubble denotes our checklist-guided MLLM evaluation, ArtifactsBench, which achieves 94.4% agreement with human… See the full description on the dataset page: https://huggingface.co/datasets/tencent/ArtifactsBenchmark.