Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
GRPO_LLAMA3-instructive_reasoning1 – AI Model by alibidaran | AlphaNeural AI
You can deploy this model and start earning money today!
alibidaran
/
GRPO_LLAMA3-instructive_reasoning1
like
0
transformers
safetensors
text-generation-inference
unsloth
llama
trl
en
apache-2.0
endpoints_compatible
us
Views
No views yet
Model card
Files and Versions
Community
API
Deploy
Uploaded model
Developed by:
alibidaran
License:
apache-2.0
Finetuned from model :
unsloth/meta-llama-3.1-8b-instruct-unsloth-bnb-4bit
This llama model was trained 2x faster with
Unsloth
and Huggingface's TRL library.
Evalution results
We are using MMLU dataset in different tasks. Here are the results of using 100 random samples of MMLU dataset.
Professional Psychology : 76%
Manengment: 74%
sociology: 75%