Learning after pre-training
Mid-Training, SFT, reinforcement learning, and on-policy distillation for capable, reliable foundation models.
UCAS / CASIA / BEIJING
Master’s student · Class of 2028
I am a master’s student in Pattern Recognition and Intelligent Systems at the University of Chinese Academy of Sciences and the Institute of Automation, Chinese Academy of Sciences.
My research focuses on LLM mid/post-training, agents, and multimodal reinforcement learning. I am interested in how models learn to reason, use tools, and connect understanding with generation.
I am currently a research intern at Meituan LongCat Interaction through the Beidou Talent Program, following research internships at Alibaba Amap and vivo.
As of Sep. 2026
Mid-Training, SFT, reinforcement learning, and on-policy distillation for capable, reliable foundation models.
Tool use, long-horizon interaction, and efficient learning from multi-turn agent experience.
Unified multimodal evaluation, video understanding, visual reasoning, and preference alignment.
A selection of my work on learning, alignment, and evaluation. First / co-first-author work is highlighted.
7 publications
Can a model recover the prompt behind a generated video? A closed-loop benchmark with 900 verified videos and 18 VLMs connects video understanding to generation control.
Evaluating understanding and generation together: 4,104 examples across 30 subtasks reveal the gap between strong individual capabilities and truly unified multimodal reasoning.
Making visual reinforcement learning more reliable with uncertainty-aware optimization and perceptual supervision. KADID PLCC / SRCC improves from 0.723 / 0.719 to 0.779 / 0.775 over VisualQuality-R1.
Multimodal preference alignment with 120K human-annotated preference pairs, critique-based reward modeling, and dynamic preference optimization, evaluated on 27 benchmarks.
Beyond reading a single frame: 2,000 human-annotated questions probe text recognition, cross-frame integration, and spatiotemporal reasoning in video.
Learning shared and view-specific representations under missing views and labels. Evaluated against 10 methods on five datasets, with stronger average precision under 50% missingness.
Turning incomplete supervision into reliable learning signals through uncertainty-aware pseudo-labeling and label-guided dual graph constraints, validated on six datasets.
Y. Yang, H. Tian, Y. Shi, Wulin Xie, et al. · TPAMI · Under Review
J. Chen, Wulin Xie, et al. · Knowledge-Based Systems
X. Yang, Wulin Xie, et al. · Pattern Recognition
Beidou Talent Program · Research Intern

Algorithm Research Intern

Algorithm Research Intern

Research Intern

Research Intern
Institute of Automation, Chinese Academy of Sciences
M.Sc. in Pattern Recognition and Intelligent Systems · Expected Jul. 2028
B.Sc. in Information Management and Information Systems
GPA: 4.25 / 5.0
LET’S CONNECT
I’m looking for research internships in foundation models, agents, and multimodal learning. If our interests overlap, I’d love to hear from you.
xwl1085930920@gmail.com