Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
a3-rl-SankalpKJ_swesmith-oracle-filtered-40-8B – AI Model by penfever | AlphaNeural AI
You can deploy this model and start earning money today!
penfever
/
a3-rl-SankalpKJ_swesmith-oracle-filtered-40-8B
like
0
safetensors
qwen3
reinforcement-learning
rl
swe
agent
skyrl
SankalpKJ/swesmith-oracle-filtered
laion/GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink
finetune
apache-2.0
us
Views
No views yet
Model card
Files and Versions
Community
API
Deploy
a3-rl-SankalpKJ_swesmith-oracle-filtered-40-8B
RL-tuned (SkyRL) agentic SWE model. This is the
global_step 40
checkpoint, selected as the best checkpoint by EMA reward (EMA reward 0.0877).
Base model:
laion/GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink
(a Qwen3-8B SFT)
Training dataset:
SankalpKJ/swesmith-oracle-filtered
Architecture:
Qwen3-8B
Training framework:
SkyRL (agentic RL, a3 series)
The RL training configuration is included in this repo as
rl_config.yaml
. Parsed training metrics and the raw training logs are in
training_logs/
.
Training Traces
The full agentic rollout traces from this RL run are published as a dataset:
Traces:
open-athena/a3-rl-SankalpKJ_swesmith-oracle-filtered
Note
Published to the
penfever/
namespace pending a
laion/
write-role bump (may be re-homed to
laion/
later).