The
AmirhoseinGH/DS-Qwen-7b-GG-CalibratedConfRL model is derived from the
DeepSeek-R1 Qwen Distill 7B base model, optimized through
confidence-based Reinforcement Learning using GRPO for enhanced intrinsic confidence calibration.
It relates to the paper
Guided by Gut: Efficient Test-Time Scaling with Reinforced Intrinsic Confidence.
This calibration process significantly improves the reliability of the model’s internal confidence signals. The model is optimized for use with the Guided by Gut (GG) framework, a self-guided test-time scaling (TTS) strategy that leverages these intrinsic confidence signals to perform complex reasoning tasks efficiently—without costly external verifier models.
Traditional TTS methods often require substantial computational resources due to their reliance on external verifier models like Process Reward Models (PRMs) or extensive sampling strategies (e.g., Best-of-N). The GG framework provides a powerful yet computationally efficient alternative: