The Arabic Broad Benchmark is a unique dataset and an advanced benchmark created by SILMA.AI to assess the performance of Large Language Models in Arabic Language.
ABB consists of 470 high quality human-validated questions sampled from 64 Arabic benchmarking datasets, evaluating 22 categories and skills.
The advanced benchmarking script utilizes the dataset to evaluate models or APIs using a mix of 20+ Manual Rules and LLM as Judge variations customized… See the full description on the dataset page:
https://huggingface.co/datasets/silma-ai/arabic-broad-benchmark.