NVIDIA
Current May 2026 – PresentApplied Deep Learning Research Intern · Mentor: Wei Ping
Token-efficient reinforcement learning for LLM reasoning — getting more capability per training and inference token.
I am a Ph.D. student in Machine Learning at Georgia Institute of Technology, advised by Prof. Tuo Zhao. I received my bachelor's and master's degrees from the University of Science and Technology of China. My research focuses on efficient and reliable large language model training, including optimizer and architecture design for pretraining, as well as training algorithms for mid- and post-training.
Recent milestones, newest first.
Working on token-efficient reinforcement learning for LLM reasoning with Wei Ping.
Our optimizer set new records on the modded-nanogpt speedrun leaderboard and was adopted in Andrej Karpathy's nanochat.
The work led to NorMuon and Shuffle the Context (long-context adaptation via RoPE-perturbed self-distillation), both published at ICML 2026.
SlimMoE (COLM 2025) shipped as Phi-mini-MoE and Phi-tiny-MoE on Hugging Face — 6M+ total downloads.
Adaptive preference scaling for preference optimization, and robust reward learning from corrupted human feedback.
Research internships at industry labs.
Applied Deep Learning Research Intern · Mentor: Wei Ping
Token-efficient reinforcement learning for LLM reasoning — getting more capability per training and inference token.
Research Intern · Mentor: Chen Liang
Long-context adaptation of LLMs — RoPE-perturbed self-distillation, published as Shuffle the Context at ICML 2026.
Research Intern · Mentor: Chen Liang
Structured compression and distillation of large MoE models — shipped as Phi-mini-MoE and Phi-tiny-MoE (6M+ downloads), published as SlimMoE at COLM 2025.
* denotes equal contribution.
New records on the modded-nanogpt speedrun; adopted in Andrej Karpathy's nanochat.
Released as Microsoft's open Phi-mini-MoE and Phi-tiny-MoE — 6M+ downloads.
A standard method for neural temporal point processes — co-developed as an undergraduate.
Ph.D. in Machine Learning
Advisor: Prof. Tuo Zhao
M.S. in Data Science
Advisor: Prof. Lan Zhang
B.S. in Mathematics and Applied Mathematics · School of the Gifted Young
Entered university at 14 through USTC's selective early-entrance program; graduated at 18.