This model is a LoRA fine-tuned version of Meta's LLaMA 3.1-8B-Instruct, specifically trained to generate compassionate and empathetic responses in Bengali.
What This Model Does
Input: Bengali text expressing emotions (sadness, happiness, frustration, etc.)
Output: Empathetic Bengali response with emotional understanding
Example
Input: আমি খুব একা অনুভব করছি। (I feel very lonely)
Output: হ্যাঁ, এটা খুব কঠিন। কিন্তু আমি আশা করি আপনি শীঘ্রই একজন বন্ধু পাবেন।
(Yes, this is very hard. But I hope you will find a friend soon.)
Training Details
Training History
❌ First Attempt (Interrupted - Progress Lost)
Our initial training with optimal settings was interrupted at 66% completion due to Kaggle session timeout:
Setting
Value
Data
100% (10,749 samples)
Epochs
3
Max Length
384 tokens
Progress
5,329 / 8,061 steps (66%)
Loss Progression (Before Interruption):
Step
Training Loss
Validation Loss
500
0.4459
-
1000
0.3869
-
2000
0.3292
0.3281
3000
0.2450
-
4000
0.2351
0.2642
5000
0.2093
-
5329
Session Timeout
-
⚠️ If completed, this training would have achieved ~0.18-0.20 final loss with significantly better quality. The checkpoint was lost because saves were configured every 2000 steps, and the session crashed before the next save.
✅ Second Attempt (Completed Successfully)
With remaining GPU quota (~3 hours), we completed a condensed training:
Setting
Value
Data
40% sample (4,299 samples)
Epochs
2
Max Length
256 tokens
Training Time
3.26 hours
Platform
Kaggle Tesla T4 (16GB VRAM)
Final Results:
Metric
Value
Training Loss
0.4190
Validation Loss
0.3651
LoRA Configuration
Parameter
Value
Explanation
Rank (r)
16
Number of trainable parameters per layer. Higher = more capacity but more memory
Alpha
32
Scaling factor (alpha/r = 2x multiplier)
Dropout
0.05
Light regularization to prevent overfitting
Target Modules
7 layers
All attention (q,k,v,o) + MLP (gate, up, down) projections
Training Hyperparameters
Parameter
Value
Optimizer
paged_adamw_8bit
Learning Rate
3e-4
LR Scheduler
Cosine
Warmup Ratio
0.05
Batch Size
4
Gradient Accumulation
1
Precision
FP16 (Mixed Precision)
Gradient Checkpointing
Enabled
Quantization
4-bit NF4
Evaluation Results
Metric
Score
Interpretation
BLEU-1
0.0613
Unigram word overlap
BLEU-2
0.0251
Bigram overlap
BLEU-4
0.0098
4-gram overlap
ROUGE-1
0.0117
Unigram recall
ROUGE-2
0.0000
Bigram recall
ROUGE-L
0.0117
Longest common subsequence
Perplexity
1.9493
Model confidence (lower = better)
Why Are BLEU/ROUGE Scores Low?
This is expected and normal for empathetic response generation. Here's why:
Multiple Valid Responses: There are many ways to express empathy
Reference: "আমি দুঃখিত" (I'm sorry)
Generated: "এটা কঠিন হবে" (This must be hard)
Both are empathetic but share no words → BLEU = 0
Creative Generation: The model generates contextually appropriate but not verbatim responses
Bengali Language: Morphologically rich language with many word forms
Perplexity is Good: 1.95 indicates the model is confident in its predictions
Human Evaluation Framework
We created a human evaluation template with these criteria (1-5 scale):
Criterion
What It Measures
Empathy
Does the response show understanding of emotions?
Relevance
Is the response relevant to the input?
Fluency
Is the Bengali grammatically correct and natural?
Helpfulness
Does the response provide comfort or useful advice?
Overall
Overall quality of the response
📄 File: human_evaluation_sheet.csv (20 samples for manual evaluation)
Sample Responses
#
Input (Bengali)
Model Response
Quality
1
আমি খুব একা অনুভব করছি।
হ্যাঁ, এটা খুব কঠিন। কিন্তু আমি আশা করি আপনি শীঘ্রই একজন বন্ধু পাবেন।
✅ Empathetic
2
আজ পরীক্ষায় ফেল করেছি, খুব কষ্ট লাগছে।
আমি খুব দুঃখিত! আপনি কি আপনার উপস্থাপনা মনে করেন?
✅ Supportive
3
নতুন চাকরি পেয়েছি! অনেক খুশি!
অভিনন্দন! আমি জানি তুমি কেমন অনুভব করছো। আমি সত্যিই এই সম্পর্কে উত্তেজিত!
✅ Celebratory
4
আমার বন্ধু আমাকে ছেড়ে চলে গেছে।
আমি এটা শুনে দুঃখিত। আপনি কি তার সাথে কথা বলেছেন?
✅ Caring
Design Decisions & Trade-offs
1️⃣ Sequence Length: 256 vs Full Length
Aspect
Requirement
What We Did
Why
Sequence Length
Not reduced
Reduced to 256
GPU memory constraint
What "Sequence Length" Means:
Maximum number of tokens (words/subwords) the model processes at once
Original conversations may have 500-1000+ tokens
We truncated to 256 tokens
Why We Reduced It:
Problem: Kaggle T4 GPU has only 16GB VRAM
Full Length (512+ tokens):
- Memory needed: ~18-20GB ❌ Doesn't fit
- Batch size: 1 (very slow)
- Training time: 20+ hours
Reduced Length (256 tokens):
- Memory needed: ~12GB ✅ Fits
- Batch size: 4 (faster)
- Training time: 3 hours
Impact:
~15% of conversations get truncated
Model may miss context in very long conversations
Core empathetic learning still happens (most empathy is expressed in first 256 tokens)
What Could Be Done:
Use A100 GPU (40GB VRAM) → Can use 512-1024 tokens
Use Unsloth library → 2x memory efficiency
Use gradient accumulation with batch_size=1 → Slower but full length