google/flan-t5-large adapted for dialogue summarization on
DialogSum, trained as part of an
IIT-D Gen-AI course project comparing four fine-tuning methods under identical conditions.
Code, evaluation harness and the other three models:
https://github.com/dipika-s/iitd-genai
1from transformers import AutoModelForSeq2SeqLM, AutoTokenizer
2from peft import PeftModel
3
4base = AutoModelForSeq2SeqLM.from_pretrained("google/flan-t5-large")
5model = PeftModel.from_pretrained(base, "daggar/flan-t5-dialogsum-qlora")
6tokenizer = AutoTokenizer.from_pretrained("daggar/flan-t5-dialogsum-qlora")
7
8dialogue = "#Person1#: Hi, how was your weekend?\n#Person2#: Great, I went hiking."
9inputs = tokenizer("Summarize the following dialogue:\n" + dialogue,
10 return_tensors="pt", max_length=512, truncation=True)
11print(tokenizer.decode(model.generate(**inputs, max_new_tokens=128)[0],
12 skip_special_tokens=True))
Measured on the full 1,500-example DialogSum test split, against the untuned base model.
* see the known issue on the prefix model card.
Trained only on DialogSum, which is two-speaker English conversation transcripts using
#Person1# / #Person2# speaker tags. Summaries of longer, multi-party, domain-specific
or non-English dialogue will be unreliable. Dialogues over 512 tokens are truncated, so
content late in a long conversation may be dropped. The model inherits any biases present
in google/flan-t5-large and in DialogSum, and summaries can contain details not supported by the
source dialogue.