Saliency Guided Reinforcement Leaming for Multi Step Mathematical Reasoning in Large Language Models
Date of Award
2026
Document Type
Thesis
Publisher
Santa Clara : Santa Clara University, 2026
Degree Name
Master of Science (MS)
Department
Computer Science and Engineering
First Advisor
Yi Fang
Abstract
Large language models perform strongly on many tasks but remain unreliable on complex multi step mathematical reasoning. Existing reinforcement learning methods treat reasoning trajectories as uniformly informative, ignoring that some intermediate steps matter more than others. This thesis argues that reasoning is structurally uneven and that optimization should reflect this internal credit distribution.
We adopt Group Relative Policy Optimization (GRPO) to construct a fully automated reinforcement learning framework. To provide intermediate supervision without human annotation, we automatically generate short quiz questions that test key reasoning checkpoints. Quiz rewards complement final answer rewards, enabling structured feedback while keeping the pipeline scalable. Supervised fine tuning is explored for initialization, but reinforcement learning drives the main gains.
The central contribution is a saliency guided token reweighting mechanism.We perform chunk level saliency analysis to estimate the association between each reasoning chunk and final correctness. The analysis shows that only a subset of reasoning chunks carries decisive influence. We therefore selectively reweight token level losses in GRPO, amplifying weak reasoning chunks only when their advantage is positive, in order to strengthen constructive credit assignment during optimization.
Across GSM8K, MATH, MATH 500, and AIME benchmarks, saliency guided GRPO yields consistent and often superior performance compared to supervised baselines and standard GRPO variants. More complex reward designs yield limited additional benefit, suggesting that structurally informed credit allocation is more effective than simply adding more reward components.
Recommended Citation
Lo, Ji Dung, "Saliency Guided Reinforcement Leaming for Multi Step Mathematical Reasoning in Large Language Models" (2026). Computer Science and Engineering Master's Theses. 62.
https://scholarcommons.scu.edu/cseng_mstr/62
