Model Card for t5_small Summarization Model
Model Details
This model is a fine-tuned version of T5-small for text summarization tasks using the CNN/DailyMail dataset.
Training Data
The model was trained on a subset (1%) of the CNN/DailyMail dataset, which consists of news articles and their corresponding highlights.
Training Procedure
- Learning Rate: 2e-5
- Epochs: 1
- Batch Size: 4
- Max Length: 512
How to Use
from transformers import AutoTokenizer, AutoModelForSeq2SeqLM
tokenizer = AutoTokenizer.from_pretrained("./latest_checkpoint")
model = AutoModelForSeq2SeqLM.from_pretrained("./latest_checkpoint")
Evaluation
Loss: 0.211
ROUGE-1: 1.59
ROUGE-2: 0.66
ROUGE-L: 1.39
BLEU-1: 61.39
BLEU-2: 30.85
BLEU-4: 11.25
Limitations
The model may occasionally omit important details or introduce factual inconsistencies in the generated summaries. It also has limited understanding of context in very long articles.
Ethical Considerations
Bias: The model may reflect biases present in the CNN/DailyMail dataset.
Factual Accuracy: Users should verify the accuracy of generated summaries before use, especially in critical applications.