I am a master's student at Peking University. I received my
undergraduate degree from
Beijing University of Posts and Telecommunications.
My research focuses on large language models and multimodal understanding. I am
especially interested in post-training methods, including rejection sampling
fine-tuning and reinforcement learning, as well as self-evolution mechanisms for
large models. I also work on multimodal large models for scene understanding and
recognition.
Research Interests
LLM Post-trainingRejection Sampling Fine-tuningReinforcement LearningSelf-evolution of LLMsMultimodal Large ModelsScene Graph Generation
Education
PKU
Peking University
Master's Student
Beijing, China
BUPT
Beijing University of Posts and Telecommunications
Undergraduate Student
Beijing, China
News
One paper has been accepted to ACM Multimedia 2026.
A recent work on eliciting native reasoning in base models through
small-data self-distillation is currently under submission.
SGG-R³ has been accepted to ACL Findings 2026.
Exploring the self-evolution of reasoning in large models and probing their
reasoning boundaries through reinforcement learning and rejection sampling
fine-tuning.