Current
Qwen Team, Alibaba Group
Researcher
Model Architecture & Optimization
I study the architectures and training methods behind large language models.
My work focuses on model architecture and optimization, including attention mechanisms, mixture-of-experts models, and efficient pre-training.
arXiv
Architecture, efficiency, and training stability in sparse language models.
CoLM 2026
A unified account of attention and residual sinks, linking outlier-driven rescaling to training stability. * Equal contribution.
NeurIPS 2025 / Best Paper Award
Understanding how gating improves attention, model performance, and training stability. * Equal contribution.
ACL 2025
Global-batch load balancing for expert specialization, used in Qwen3-MoE models. * Equal contribution.
Technical report
Dense and mixture-of-experts language models with unified thinking and non-thinking modes.
Technical report
The Qwen2.5 family of pretrained and instruction-tuned large language models.
Current
Researcher
Model Architecture & Optimization
2020–2021 · 2018–2019
Research Intern
Mentors: Li Dong (2020–2021);
Nan Duan (2018–2019)
Ph.D. in Computer Science and Technology
HIT-SCIR · Advisor: Prof. Wanxiang Che
2016–2018 · 2012–2016
M.S. & B.S. in Computer Science and Technology
2025
Gated Attention for Large Language Models · Co-first author
2022
EMNLP workshop · HIT-SCIR team · 1st place in the full-dataset track
2013–2015
2 gold and 4 silver medals, including one regional runner-up finish
2016
China Computer Federation
2015, 2017, 2021
Three-time recipient