Enhancing Logical Reasoning in Language Models via Symbolically-Guided Monte Carlo Process Supervision (sft)
γ π¦ GitHub repo | π€ Paper γ
TL;DR
We propose to enhance LLMs' logical reasoning and generalization by synthesizing symbolic reasoning trajectories using Monte Carlo estimation and integrating them with Direct Preference Optimization and Supervised Fine-Tuning.