Paper | Code Repository
LiveEvalBench is a benchmark for evaluating LLM-generated frontend code (Vue / React / static HTML). Each evaluation actually runs the generated project locally and drives a multi-role agent panel against the live app. This repository hosts the benchmark dataset (the queries and their grading rubrics); the evaluation framework code lives in the LiveEvalBench code repository.
100 user requests (queries) for… See the full description on the dataset page:
https://huggingface.co/datasets/wyysteelhead/LiveEvalBench.