Biography

I am a master's student at Peking University. I received my undergraduate degree from Beijing University of Posts and Telecommunications.

My research focuses on large language models and multimodal understanding. I am especially interested in post-training methods, including rejection sampling fine-tuning and reinforcement learning, as well as self-evolution mechanisms for large models. I also work on multimodal large models for scene understanding and recognition.

Research Interests

LLM Post-training Rejection Sampling Fine-tuning Reinforcement Learning Self-evolution of LLMs Multimodal Large Models Scene Graph Generation

Education

PKU

Peking University

Master's Student

Beijing, China

BUPT

Beijing University of Posts and Telecommunications

Undergraduate Student

Beijing, China

News

  • One paper has been accepted to ACM Multimedia 2026.
  • A recent work on eliciting native reasoning in base models through small-data self-distillation is currently under submission.
  • SGG-R³ has been accepted to ACL Findings 2026.
  • Exploring the self-evolution of reasoning in large models and probing their reasoning boundaries through reinforcement learning and rejection sampling fine-tuning.

Selected Publications

More on GitHub
Overview figure for SGG-R3 relation augmentation and two-stage post-training

SGG-R³: From Next-Token Prediction to End-to-End Unbiased Scene Graph Generation

Findings of ACL, 2026

A study on end-to-end unbiased scene graph generation, connecting multimodal perception with structured visual reasoning.

Projects

Foundations of RL Learning

A hands-on educational library with from-scratch Python implementations of foundational reinforcement learning algorithms.