Authors
Zhenwen Liang, Sidi Lu, Wenhao Yu, Kishan Panaganti, Yujun Zhou, Haitao Mi, and Dong Yu
Context
This work grew out of a close collaboration with Zhenwen Liang during my time at Tencent AI Lab Seattle.
Overview
G2RL is a gradient-guided reinforcement-learning framework for language-model reasoning. Instead of using token entropy or an external embedding model as a proxy for exploration, it asks whether a sampled trajectory would move the policy in a genuinely different direction.
The method derives a sequence-level feature from the model’s final-layer sensitivity, compares these features within each rollout group, and uses the result as a bounded reward multiplier. Trajectories that add useful gradient directions receive more weight, while redundant or off-manifold updates are deemphasized.
Results
We evaluated G2RL on Qwen3 base models at 1.7B and 4B parameters across mathematical and general-reasoning benchmarks, including MATH500, AMC, AIME24, AIME25, GPQA, and MMLU-Pro. It consistently improved pass@1, majority-vote accuracy, and pass@k over entropy-based GRPO and external-embedding baselines.
Paper
Status
Preprint, 2025.
Citation
Zhenwen Liang, Sidi Lu, Wenhao Yu, Kishan Panaganti, Yujun Zhou, Haitao Mi, and Dong Yu. “Can LLMs Guide Their Own Exploration? Gradient-Guided Reinforcement Learning for LLM Reasoning.” arXiv:2512.15687, 2025.