A crypto sentiment analysis model fine-tuned on Qwen3-8B using SFT + DPO (Direct Preference Optimization).
Model Description
Senti is a specialized LLM for cryptocurrency market sentiment analysis. It generates structured market summaries from social media data, classifies sentiment, and answers questions about crypto markets.
Training Pipeline
Qwen/Qwen3-8B (Base, 8B params)
│
▼
┌─────────────────────────────────────┐
│ SFT Training │
│ 9,314 samples (train + val) │
│ 5.34 hours on A100 80GB │
└─────────────────────────────────────┘
│
▼
┌─────────────────────────────────────┐
│ DPO Training │
│ 1,477 preference pairs │
│ 30.7 min, 99.4% accuracy │
└─────────────────────────────────────┘
│
▼
This Model
Training Data
Overview
Dataset
Samples
Size
Purpose
SFT Training
8,383
22.4 MB
Supervised fine-tuning
SFT Validation
931
2.4 MB
Evaluation
DPO Preference
1,477 pairs
4.8 MB
Preference alignment
Total
10,791
29.6 MB
SFT Dataset Details
Task Distribution
Task
Samples
Percentage
Description
Summary Generation
4,491
53.6%
Generate market summaries from tweets
Sentiment Classification
3,624
43.2%
Classify individual tweet sentiment
Chat / Q&A
268
3.2%
Answer crypto market questions
Summary Task Breakdown
Summary Type
Count
Description
Hourly Summaries
4,086
Real-time market pulse from recent tweets
Daily Summaries
405
End-of-day market recap and synthesis
Tweets Per Summary
Metric
Value
Average
44.4 tweets
Minimum
5 tweets
Maximum
100 tweets
Total Tweets Processed
~199,000
Input/Output Statistics
Metric
Characters
~Tokens
Avg Input Length
2,167
~540
Avg Output Length
995
~250
Max Sequence Length
4,096
4,096
Sentiment Classification Details
5-Class System:
Label
Description
Example
STRONG_BULLISH
Extreme optimism, price targets, euphoria
"BTC to 200k! This is just the beginning! 🚀🚀🚀"
BULLISH
Positive outlook, accumulation mentions
"Adding more ETH here, looking good for Q1"
NEUTRAL
Factual, news, no clear sentiment
"Bitcoin trading at $97,500 today"
BEARISH
Concern, caution, negative outlook
"Not liking this price action, reducing exposure"
STRONG_BEARISH
Panic, crash predictions, extreme fear
"This is the top, selling everything, bear market incoming"
Classification Distribution (3,624 samples):
Label
Count
Percentage
BULLISH
~1,200
33.1%
NEUTRAL
~1,100
30.3%
STRONG_BULLISH
~650
17.9%
BEARISH
~450
12.4%
STRONG_BEARISH
~224
6.2%
Coins Covered
The model was trained on data from 12 cryptocurrencies:
Coin
Ticker
Category
Sample Distribution
Bitcoin
BTC
Major
~18%
Ethereum
ETH
Major
~16%
Solana
SOL
Layer 1
~12%
XRP
XRP
Payment
~10%
BNB Chain
BNB
Layer 1
~9%
Chainlink
LINK
Oracle
~8%
Arbitrum
ARB
Layer 2
~7%
Avalanche
AVAX
Layer 1
~6%
Hyperliquid
HYPE
DeFi
~5%
0G
0G
AI/Infra
~4%
Aster
ASTER
Emerging
~3%
Zcash
ZEC
Privacy
~2%
Data Sources
Source
Description
Twitter/X
Real-time tweets from crypto KOLs and communities
Time Range
October 2025 - January 2026
Languages
English (primary), Turkish (secondary)
Filtering
Relevance scoring, spam removal, deduplication
DPO Preference Dataset Details
Overview
DPO (Direct Preference Optimization) aligns the model to produce higher-quality, more detailed responses by learning from preference pairs.
Property
Value
Total Pairs
1,477
Format
(prompt, chosen, rejected) triplets
Purpose
Improve response quality and structure
Task Distribution
Task
Pairs
Percentage
Description
Summary Generation
877
59.4%
Prefer detailed market analysis
Chat / Q&A
600
40.6%
Prefer informative, grounded answers
Language Distribution
Language
Pairs
Percentage
English
1,193
80.8%
Turkish
284
19.2%
Response Length Statistics
Metric
Chosen
Rejected
Average Length
1,836 chars
249 chars
Minimum Length
1,026 chars
198 chars
Maximum Length
2,719 chars
282 chars
Avg Ratio
7.5x longer
baseline
Length Ratio Distribution
Ratio Range
Description
Minimum
4.0x (chosen vs rejected)
Average
7.5x
Maximum
13.7x
Preference Pair Generation Methodology
Component
Method
Chosen Responses
Generated by GPT-4o with detailed prompts
Rejected Responses
Systematically degraded versions
Chosen Response Characteristics:
Well-structured with markdown headers
Specific data points (prices, percentages, names)
Actionable insights and analysis
Consistent formatting across responses
Appropriate length (1,000-2,700 chars)
Rejected Response Degradation Methods:
Truncation: Cut to 1-2 generic sentences
Data Removal: Strip specific numbers, names, prices
Generic Templates: "The market is bullish. Prices may rise."
Structure Removal: No headers, bullets, or formatting
Vague Language: Replace analysis with platitudes
Example Preference Pair:
Prompt: "Summarize BTC market sentiment from these tweets: [100 tweets about Bitcoin ETF inflows]"
Chosen (1,850 chars):
"## BTC Hourly Summary
**Overall Sentiment: BULLISH (7.8/10)**
### Key Developments
- Bitcoin broke through $98,000 resistance with strong volume
- ETF inflows reached $500M for the third consecutive day
- Whale accumulation: 3 wallets acquired 2,500+ BTC in 24h
### Risk Factors
- RSI approaching overbought (72)
- Some profit-taking expected at $100K psychological level
### Outlook
Momentum strong. Short-term target: $102,000. Support at $95,500."
Rejected (198 chars):
"Bitcoin sentiment is positive today. The market shows bullish signs.
Traders are optimistic about price movements. Consider doing your own research."
Input: Raw tweets about a cryptocurrency
Output: Structured market analysis
markdown
1## BTC Hourly Summary23**Overall Sentiment: BULLISH (7.8/10)**45### Key Developments6- Bitcoin broke through $98,000 resistance with strong volume
7- Whale accumulation detected: 3 wallets acquired 2,500+ BTC
8- ETF inflows reached $500M for the third consecutive day
910### Risk Factors11- RSI approaching overbought territory (72)
12- Some profit-taking expected at $100K psychological level
1314### Outlook15Momentum remains strong. Short-term target: $102,000.
2. Sentiment Classification
Input: Single tweet
Output: Sentiment label + confidence
Input: "BTC breaking 100k! This is the moment we've been waiting for! 🚀"
Output: STRONG_BULLISH (confidence: 0.94)
Supported Labels:
STRONG_BULLISH
BULLISH
NEUTRAL
BEARISH
STRONG_BEARISH
3. Chat / Q&A
Input: Question about crypto market
Output: Data-grounded answer
Q: "What's driving ETH sentiment today?"
A: "Based on recent data, ETH sentiment is moderately bullish (6.5/10).
Key drivers include:
1. Successful Dencun upgrade mentions (+45% positive)
2. Layer 2 TVL growth reaching $15B
3. Some concerns about gas fees during high activity periods
The overall trend suggests cautious optimism."
Usage
Basic Usage
python
1from peft import PeftModel
2from transformers import AutoModelForCausalLM, AutoTokenizer
3import torch
45# Load model6base_model = AutoModelForCausalLM.from_pretrained(7"Qwen/Qwen3-8B",8 torch_dtype=torch.bfloat16,9 device_map="auto"10)11model = PeftModel.from_pretrained(base_model,"dorukardahan/senti-qwen3-8b-dpo")12tokenizer = AutoTokenizer.from_pretrained("Qwen/Qwen3-8B")1314# Generate15messages =[16{"role":"system","content":"You are Senti, a crypto sentiment analyst."},17{"role":"user","content":"Analyze BTC market sentiment based on recent social media activity."}18]1920text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)21inputs = tokenizer(text, return_tensors="pt").to(model.device)2223outputs = model.generate(24**inputs,25 max_new_tokens=512,26 temperature=0.7,27 do_sample=True28)29print(tokenizer.decode(outputs[0], skip_special_tokens=True))