Large Language Models · Reasoning · RLVR

Long Li

I previously worked with Prof. Chao Qu and Prof. Yuan Qi at Fudan University on reinforcement learning for large language models, with multiple first-author and co-first-author publications at ICLR, ICML, ACL, and EMNLP. I am now a Ph.D. student at Griffith University advised by Prof. Shirui Pan, whose work has received 60,000+ citations, and I work on world models and embodied intelligence.

Portrait of Long Li
Ph.D. student at Griffith University; M.S. in Computer Science, East China Normal University.

Research

Research Interests

RL & RLVR for LLMs

Designing reinforcement learning methods for LLM reasoning, especially RLVR settings where verifiable rewards can improve correctness while preserving exploration and diversity.

Agentic RL

Training agentic systems with dense feedback, trajectory aggregation, and multi-turn decision making, including SQL agents and broader tool-using LLM agents.

Emerging Directions

Exploring multi-agent reinforcement learning, agentic RL frameworks, and Vision-Language-Action models for embodied and interactive intelligence.

Publications

Selected Publications

2026
ICML 2026 · Co-first Author

DisPPO: Quantile-Based Distributional Reinforcement Learning for Large Language Models

Zhijian Zhou*, Long Li*, Xuan Zhang, Zongkai Liu, Yanting Miao, Yuchen Liu, Deshu Chen, Ke Li, Xing Sun, Ruoxi Jiang, Xiaoyu Tan, Chao Qu, Yuan Qi.

Develops a quantile-based distributional reinforcement learning method for LLM training, improving policy optimization by modeling the return distribution rather than only its expectation.

2026
ICML 2026 · Co-first Author

Anchored Policy Optimization: Mitigating Exploration Collapse Via Support-Constrained Rectification

Tianyi Wang*, Long Li*, Hongcan Guo, Yibiao Chen, Yixia Li, Yong Wang, Yun Chen, Guanhua Chen.

Introduces Anchored Policy Optimization (APO), a support-constrained RLVR method that mitigates exploration collapse and improves the accuracy-diversity trade-off in mathematical reasoning.

2026
ICLR 2026 · First Author

The Choice of Divergence: A Neglected Key to Mitigating Diversity Collapse in Reinforcement Learning with Verifiable Reward

Long Li, Jiaran Hao, Jason Klein Liu, Zhijian Zhou, Xiaoyu Tan, Wei Chu, Zhe Wang, Shirui Pan, Chao Qu, Yuan Qi.

Proposes replacing reverse KL with f-divergence based sampling from a reference model, improving efficiency and entropy behavior across SQL and math tasks for 7B-32B models.

2026
ACL Findings 2026 · First Author

SQL-ASTRA: Alleviating Sparse Feedback in Agentic SQL via Column-Set Matching and Trajectory Aggregation

Long Li, Zhijian Zhou, Jiangxuan Long, Peiyang Liu, Weidi Xu, Zhe Wang, Shirui Pan, Chao Qu.

Builds a multi-turn RL framework for SQL agents with per-turn dense rewards and aggregated intermediate supervision.

2025
ACL Findings 2025 · Co-first Author

To Code or not to Code? Adaptive Tool Integration for Math Language Models via Expectation-Maximization

Haozhe Wang*, Long Li*, Chao Qu, Weidi Xu, Fengming Zhu, Yi Xin, Wei Chu, Fangzhen Lin.

2024
EMNLP 2024 · First Author

How Do Humans Write Code? Large Models Do It the Same Way Too

Long Li, Xuzheng He, Haozhe Wang, Linlin Wang, Liang He.

Proceedings of EMNLP 2024, pages 4638-4649, Miami, Florida, USA.

Experience

Industry Research

I work on practical LLM training systems, from data construction and SFT to RL optimization and evaluation.

Algorithm Engineer · INF Technology

March 2025 - March 2026

  • Led NL2SQL leaderboard efforts across data collection, SFT/RL training, dense reward design, and evaluation.
  • Improved pass@k to pass@1 conversion through fine-grained reward modeling and f-divergence based RL constraints.
  • Helped produce a pure open-source model ranked first on BIRD Bench among open-source submissions at the time of submission.
  • Developed multi-turn RL capabilities on top of VERL, including off-policy data mechanisms and long-context rollout optimization.

Algorithm Intern · INF Technology

March 2024 - March 2025

  • Contributed to the release of the INF 34B pre-trained model.
  • Worked on mathematical capability enhancement across data collection, SFT, and RL optimization.

Education

Academic Background

Griffith University

Ph.D. student, 2026.02 - Present. Advisor: Prof. Shirui Pan.

East China Normal University

M.S. in Computer Science and Technology, 2022 - 2025.

Changchun University of Science and Technology

B.S. in Computer Science and Technology, 2018 - 2022. GPA: 87/100, top 5%.

Awards

Programming Contests

ACM-ICPC Asia Regional

Silver Medal, Shenyang Site, 2021.

CCPC Guilin Site

Silver Medal, 2021.

Lanqiao Cup

National First Prize in C++ and two-time Jilin Province Champion.

Team Programming Ladder

National First Prize in consecutive years.

Contact

Open to research conversations on LLM reasoning, RL training, and agentic SQL.