Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
blimp-textworld-blimp-echo-q8 – AI Model by andthattoo | AlphaNeural AI
You can deploy this model and start earning money today!
andthattoo
/
blimp-textworld-blimp-echo-q8
like
0
transformers
safetensors
qwen3
text-generation
blimp
textworld
reinforcement-learning
Qwen/Qwen3-1.7B
finetune
text-generation-inference
endpoints_compatible
us
Views
No views yet
Model card
Files and Versions
Community
API
Deploy
blimp-textworld-blimp-echo-q8
BLiMP 5-step block-memory RL with ECHO/score auxiliary losses on TextWorld q8.
This is a full-parameter RL fine-tuned checkpoint, not a LoRA adapter.
Base model:
Qwen/Qwen3-1.7B
Final held-out TextWorld q8 eval, 32 episodes:
untrained Qwen3-1.7B: success 0.375, mean steps 36.59
standard full-history RL: success 0.375, mean steps 35.375
BLiMP block-memory RL: success 0.53125, mean steps 33.25
BLiMP + ECHO/score: success 0.5, mean steps 33.71875
GitHub repo:
https://github.com/andthattoo/blimp