Views
No views yet
Qwen/Qwen3-0.6B, trained specifically for Dakota language grammar and translation tasks using compositional reward functions on non-coding tasks. This represents a test of the RL pipeline's effectiveness for complex, multi-component reward structures in linguistic domains.harleycooper/dakota1890 (v0.1.17)1@misc{dakota1890-rl-2024,
2 title={Qwen3-0.6B-Dakota-Grammar-RL: A Compositional Reward RL Test for Non-Coding Tasks},
3 author={Christian H. Cooper},
4 year={2024},
5 url={https://huggingface.co/harleycooper/Qwen3-0.6B-Dakota-Grammar-RL}
6}