RL & RLVR for LLMs
Designing reinforcement learning methods for LLM reasoning, especially RLVR settings where verifiable rewards can improve correctness while preserving exploration and diversity.
Large Language Models · Reasoning · RLVR
I previously worked with Prof. Chao Qu and Prof. Yuan Qi at Fudan University on reinforcement learning for large language models, with multiple first-author and co-first-author publications at ICLR, ICML, ACL, and EMNLP. I am now a Ph.D. student at Griffith University advised by Prof. Shirui Pan, whose work has received 60,000+ citations, and I work on world models and embodied intelligence.
Research
Designing reinforcement learning methods for LLM reasoning, especially RLVR settings where verifiable rewards can improve correctness while preserving exploration and diversity.
Training agentic systems with dense feedback, trajectory aggregation, and multi-turn decision making, including SQL agents and broader tool-using LLM agents.
Exploring multi-agent reinforcement learning, agentic RL frameworks, and Vision-Language-Action models for embodied and interactive intelligence.
Publications
Zhijian Zhou*, Long Li*, Xuan Zhang, Zongkai Liu, Yanting Miao, Yuchen Liu, Deshu Chen, Ke Li, Xing Sun, Ruoxi Jiang, Xiaoyu Tan, Chao Qu, Yuan Qi.
Develops a quantile-based distributional reinforcement learning method for LLM training, improving policy optimization by modeling the return distribution rather than only its expectation.
Tianyi Wang*, Long Li*, Hongcan Guo, Yibiao Chen, Yixia Li, Yong Wang, Yun Chen, Guanhua Chen.
Introduces Anchored Policy Optimization (APO), a support-constrained RLVR method that mitigates exploration collapse and improves the accuracy-diversity trade-off in mathematical reasoning.
Long Li, Jiaran Hao, Jason Klein Liu, Zhijian Zhou, Xiaoyu Tan, Wei Chu, Zhe Wang, Shirui Pan, Chao Qu, Yuan Qi.
Proposes replacing reverse KL with f-divergence based sampling from a reference model, improving efficiency and entropy behavior across SQL and math tasks for 7B-32B models.
Long Li, Zhijian Zhou, Jiangxuan Long, Peiyang Liu, Weidi Xu, Zhe Wang, Shirui Pan, Chao Qu.
Builds a multi-turn RL framework for SQL agents with per-turn dense rewards and aggregated intermediate supervision.
Haozhe Wang*, Long Li*, Chao Qu, Weidi Xu, Fengming Zhu, Yi Xin, Wei Chu, Fangzhen Lin.
Long Li, Xuzheng He, Haozhe Wang, Linlin Wang, Liang He.
Proceedings of EMNLP 2024, pages 4638-4649, Miami, Florida, USA.
Experience
I work on practical LLM training systems, from data construction and SFT to RL optimization and evaluation.
Education
Ph.D. student, 2026.02 - Present. Advisor: Prof. Shirui Pan.
M.S. in Computer Science and Technology, 2022 - 2025.
B.S. in Computer Science and Technology, 2018 - 2022. GPA: 87/100, top 5%.
Awards
Silver Medal, Shenyang Site, 2021.
Silver Medal, 2021.
National First Prize in C++ and two-time Jilin Province Champion.
National First Prize in consecutive years.
Contact