This dataset contains router activation patterns and expert selection data from OpenAI's GPT-OSS-20B mixture-of-experts model during text generation across diverse evaluation benchmarks.
GPT-OSS-20B is OpenAI's open-weight mixture-of-experts language model with 21B total parameters and 3.6B active parameters per token. This dataset captures the internal routing decisions made by the model's router networks when… See the full description on the dataset page:
https://huggingface.co/datasets/AmanPriyanshu/GPT-OSS-20B-MoE-expert-activations.