Qwen3.5-35B-OPD is a post-trained version of
Qwen3.5-35B-A3B, a multimodal Mixture-of-Experts model with 35B total parameters and 3B activated parameters.
The model is post-trained using the method introduced in
SimpleOPD: Simple Tokenizer-Agnostic On-Policy Distillation of Long-Context Reasoning, with
SU-01 serving as the teacher model. This post-training process substantially improves the model's reasoning performance across proof and mathematical reasoning benchmarks.
The values in parentheses indicate absolute improvements over the base model.
Qwen3.5-35B-OPD uses the same model architecture and inference interface as Qwen3.5-35B-A3B. Please refer to the
Qwen3.5-35B-A3B model card
for deployment instructions and recommended inference settings.
This model is released under the
Apache License 2.0.