It achieves the following results on the evaluation set:
Loss: 1.4839
It achieves the following ROUGE scores on the test set:
rouge1: 0.555556
rouge2: 0.230769
rougeL: 0.518519
rougeLsum: 0.518519
Quick human evaluation of summarization quality: the results are generally good, after visual inspection of the summaries generated on test set conversations. However it seems some entities/attributions are incorrect (saw an example where model confuses peoples' roles in multi-person chat)