QWERTY (QWen Experts in low Rank with Tiny cache and speculative Yeets) is a Qwen-2.5-based Mixture of Experts (MoE) model leveraging speculative decoding, with all linear modules converted to low-rank approximations.
QWERTY (QWen Experts in low Rank with Tiny cache and speculative Yeets) is a Qwen-2.5-based Mixture of Experts (MoE) model leveraging speculative decoding, with all linear modules converted to low-rank approximations.
Significant research has explored bias and fairness issues with language models (see, e.g.,
Sheng et al. (2021) and
Bender et al. (2021)). Predictions generated by the model may include disturbing and harmful stereotypes across protected classes; identity characteristics; and sensitive, social, and occupational groups.
Carbon emissions can be estimated using the
Machine Learning Impact calculator presented in
Lacoste et al. (2019).
Use the code below to get started with the model.