Domain Bench for Expert Specialty(DBES):
We conduct a comprehensive evaluation of expert routing behaviors across several mainstream MoE models, including Qwen3-30B (Instruct & Thinking), Qwen3-235B-Thinking, GLM-4.6 and DeepSeek-R1. To quantify the domain-specific expertise of these models, to validate the expertise in different domain, we establish a database from open-source dataset of seven different domain with 9 partitions from different source. This benchmark aggregates diverse… See the full description on the dataset page:
https://huggingface.co/datasets/Moe-lab/DBES.