Omni-router Transformer is a new Mixture-of-Experts (MoE) architecture that explicitly couples routing across layers using a shared router to learn strong and specialized experts. Omni-router's routing decisions appear to form consistent temporal segments and strutured usage across model depth, suggesting meaningful coordination between layers.
Please refer to the paper for details.
Model Details
Model Description
This model is a 8-expert MoE model (total 555M with 84M activate parameters) with standard Switch Transformer architecture.
Developed by: Apple Machine Learning Research
Model type: ASR
Language(s): English
License: apple-amlr
Uses
This model is a speech recognition model.
How to Get Started with the Model
Please refer to the github page for detailed usage.