Views
No views yet

Architecture note. Although the upstream lineage carries aGLM-4.7label (which refers to the teacher used for the cold-start SFT trajectories, not the student), this model is a Qwen3-8B. Itsconfig.jsonreportsmodel_type: qwen3,architectures: ["Qwen3ForCausalLM"], 36 layers, hidden size 4096, 32 attention heads / 8 KV heads, and a 40,960-token context — i.e. standard Qwen3-8B.
Qwen3ForCausalLM), 36 layers, hidden size 4096, 32 attention heads, 8 KV heads, RoPE θ = 1e6pymethods2test-large tasks); the policy rolls out against each task in a Daytona sandbox and is rewarded by the task's test verifier.swesmith-fixthink-pymethods2test_rl_config.json):advantage_estimator=rloo_n), no KL loss (use_kl_loss=false, kl_loss_coef=0.0)pymethods2test/SWE-Smith-style task distribution and may generalize unevenly to other domains.Evaluation: No verified agentic-benchmark numbers are published for this specific 8B RL checkpoint in the source artifact; evaluation results are TBD. (The flagship OpenThinkerAgent-32B card reports the project's benchmark suite for the 32B SFT line.)
@misc{openthoughts-agent,
author = {Team, OpenThoughts-Agent},
title = {{OpenThoughts-Agent: Data Recipes for Agentic Models}},
howpublished = {https://www.openthoughts.ai/blog/agent},
year = {2026}
}