Beta
Explore
Marketplace
Neural Labs
Chat
Wallet
Docs
v118-dpo-round-01 – AI Model by appvoid | AlphaNeural AI
You can deploy this model and start earning money today!
appvoid
/
v118-dpo-round-01
like
0
transformers
safetensors
lfm2
text-generation
dpo
preference-optimization
text-editing
rewriting
instruct
conversational
appvoid/v118
finetune
endpoints_compatible
us
Views
No views yet
Model card
Files and Versions
Community
API
Deploy
Graphite DPO fine-tune
Private Direct Preference Optimization fine-tune of
appvoid/v118
on
./v118_dpo_round_01.jsonl
.
Training
Method: full-model DPO with TRL
Reference model:
appvoid/v118
Epochs: 1.0
Beta: 0.1
Loss type:
sigmoid
Source fields:
instruction
,
text
,
output
,
rejected
Preference rows used: 97436
Prompt format: native tokenizer chat template when available