This model is intended for use in the
Gensyn RL Swarm, to finetune locally using peer-to-peer reinforcement learning post-training.
Once finetuned, the model can be used as normal in any workflow, for details on how to do this please refer to the
original model documentation.
For more details on the original model, please refer to the original repository
here.
This model is intended for use in the
Gensyn RL Swarm system, for details on model requirements when using outside of a swarm, refer to the original Qwen repo
here.
To deploy this model into a swarm and/or participate in the Gensyn Testnet, follow the instructions in the
RL Swarm repository, read about the
testnet, read the
RL Swarm overview, and/or read the
RL Swarm technical report.