dadastory/SummOrchestra-Qwen3-8B-GRPO-BRLP-SAMSUM is a dialogue-summarization model trained on the SamSum dataset through a two-stage pipeline combining Supervised Fine-Tuning (SFT) and GRPO + BRLP preference optimization.
This design yields a model with higher faithfulness, reduced hallucination, and stronger alignment with human preference judgments.
Performance Comparison
The model consistently outperforms standard SamSum-style summarizers in both loyalty and human preference metrics.
Barplot Comparison
Kde Faithfulness Comparison
Kde Human Preference Comparison
Training Overview
1. Supervised Fine-Tuning (SFT)
Trained on SamSum to establish strong grounding and structural summarization ability.